When I transitioned owlscript from being a tree walking interpreter to compiling to bytecode, one of the features that took much longer than others was the re-implementation of set/list comprehensions. Some of the challenges posed by their imp
The Aho-Corasick algorithm is a finite automaton based string match algorithm which is an adaptaion of another well known Finite State Machine based algorithm: The Knuth-Morris-Pratt(KMP) algorithm. While the KMP algorithm uses a DFA & failure func
Nine out of ten times when we reach for regular expressions its because we want to simply know "does this text contain this pattern?". A simple boolean expression: yes or no. Sometimes we want to know the position of the entire match, as in lexic
One of the major selling points of LR parsing is the ability to write expression grammars with a higher degree of ambiguity than would otherwise be allowed. When designing an expression grammar there are ways to encode the operators precedence and asso
It's no secret that the price we pay for using a DFA in the process of lexical analysis is the (potentially) enourmous transition tables which must be managed. There are many ways of representing transition tables. Anyone who has peaked a
-
Parsing Regular Expressions with MGCPGen
-
Determinization: Converting εNFA to DFA
-
Implementing Lexical Scoping: Resolving Variable Names with De Bruijn Indices
-
Improving Balance of Ternary Search Tries
-
Making Sense of LALR Parser Construction
-
Compiling Set Comprehensions to Bytecode: an exercise in managing abstractions
-
Fast Multi-Pattern String Searching with the Aho-Corasick Algorithm
-
Capture Groups: Tracking Regular Expression Sub Matches
-
Resolving Shift/Reduce Conflicts With Operator Precedence
-
Squeezing DFAs with Pair Compression