# Overview The {py:obj}`LP ` component converts an input query into an {py:obj}`AST `. It first tokenizes the query with {py:obj}`~bluebase.lp.lexer.LpLexer`, then parses the resulting tokens with {py:obj}`~bluebase.lp.parser.LpParser`. ```text query │ ▼ LpLexer │ list[LpToken] ▼ LpParser │ Ast ▼ AST ``` {py:obj}`~bluebase.lp.manager.LpManager` provides the entry point that connects these two stages. --- ## Tokens {py:obj}`~bluebase.lp.type.LpToken` represents a token produced by the lexer. Each token consists of two fields: * {py:obj}`tag`: the kind of token * {py:obj}`value`: the matched text For example: ```sql SELECT name FROM student; ``` is tokenized as: ```python [ LpToken('WORD', "SELECT"), LpToken('WORD', "name"), LpToken('WORD', "FROM"), LpToken('WORD', "student"), LpToken('SC', ";"), ] ``` --- ## Lexer {py:obj}`~bluebase.lp.lexer.LpLexer` converts an input query into a list of {py:obj}`~bluebase.lp.type.LpToken` objects. It combines all lexical rules into a single regular expression with named groups and scans the query from left to right. ### Token Rules | Tag | Pattern | Description | | ---------- | --------------------------- | --------------------- | | ``FLOAT`` | ``\d+\.\d+`` | Float literal | | ``INT`` | ``\d+`` | Integer literal | | ``BOOL`` | ``TRUE\|FALSE`` | Bool literal | | ``STRING`` | ``'[^']*'`` | String literal | | ``WORD`` | ``[a-zA-Z_*][a-zA-Z0-9_]*`` | Identifier or keyword | | ``COMP`` | ``=\|<>\|<=\|>=\|<\|>`` | Comparison operator | | ``COMMA`` | ``,`` | Comma | | ``LP`` | ``\(`` | Left parenthesis | | ``RP`` | ``\)`` | Right parenthesis | | ``DOT`` | ``\.`` | Dot | | ``SC`` | ``;`` | Semicolon | | ``WS`` | ``[ \t\n]+`` | Whitespace | Although ``WS`` is recognized by the lexer, whitespace tokens are discarded and are not passed to the parser. ### Invalid Tokens The lexer tracks the end position of each regular expression match. The start position of the next match must equal the current position. Any unrecognized part of the query therefore creates a gap between matches and causes {py:obj}`~bluebase.lp.error.LpInvalidTokenError` to be raised. The same check is performed after the final match to ensure that the entire query has been consumed. --- ## Parser {py:obj}`~bluebase.lp.parser.LpParser` converts a sequence of {py:obj}`~bluebase.lp.type.LpToken` objects into an {py:obj}`AST `. It is organized as a recursive-descent parser: parser methods correspond closely to nonterminals in the grammar and parse their subexpressions by calling other parser methods. The parser maintains: - {py:obj}`tokens`: the input token sequence - {py:obj}`positions`: the current position in the sequence ### Basic Operations | Method | Operation | | ----------------------------------------------------- | -------------------------------------------------------------- | | {py:obj}`~bluebase.lp.parser.LpParser.peek` | Return the current token without advancing. | | {py:obj}`~bluebase.lp.parser.LpParser.consume` | Consume the current token and advance. | | {py:obj}`~bluebase.lp.parser.LpParser.match_tag` | Match a token tag and advance only on success. | | {py:obj}`~bluebase.lp.parser.LpParser.match_keyword` | Match the value of a `WORD` token and advance only on success. | | {py:obj}`~bluebase.lp.parser.LpParser.expect_keyword` | Require and consume a specific keyword. | These operations separate token-stream handling from the implementation of individual grammar rules. ### Grammar Structure Most {py:obj}`parse_*` methods correspond directly to grammar nonterminals. For example: ```text attr : WORD | WORD DOT WORD attrs : attr | attrs COMMA attr ``` {py:obj}`~bluebase.lp.parser.LpParser.parse_attr` parses a single attribute reference. A ``WORD`` by itself denotes an attribute without a table name, while ``WORD DOT WORD`` denotes a qualified attribute with a table name. {py:obj}`~bluebase.lp.parser.LpParser.parse_attrs` parses one or more attributes separated by `COMMA` tokens by repeatedly calling {py:obj}`~bluebase.lp.parser.LpParser.parse_attr`. ### Statements {py:obj}`~bluebase.lp.parser.LpParser.parse` examines the initial ``WORD`` token and dispatches to the corresponding statement parser. An unsupported initial statement causes {py:obj}`~bluebase.lp.error.LpUnknownStatementError`. ### Parsing Example Consider: ```sql SELECT student.name FROM student WHERE student.id = 10; ``` The lexer produces: ```python [ LpToken('WORD', "SELECT"), LpToken('WORD', "student"), LpToken('DOT', "."), LpToken('WORD', "name"), LpToken('WORD', "FROM"), LpToken('WORD', "student"), LpToken('WORD', "WHERE"), LpToken('WORD', "student"), LpToken('DOT', "."), LpToken('WORD', "id"), LpToken('COMP', "="), LpToken('INT', "10"), LpToken('SC', ";"), ] ``` The parser decomposes this sequence as: ```text select_statement ├── "SELECT" ├── attrs │ └── attr ├── "FROM" ├── tables │ └── table ├── where_clause │ └── conds │ └── cond │ ├── attr │ ├── COMP │ └── hand │ └── value └── SC ``` The result is an {py:obj}`~bluebase.core.ast.AstSelect` containing the parsed attributes, tables, and conditions. --- ## Manager {py:obj}`~bluebase.lp.manager.LpManager` provides the interface to the {py:obj}`LP ` component. Its {py:obj}`~bluebase.lp.parser.LpParser.parse` method performs the complete pipeline: ```python tokens = lexer.tokenize(query) parser = LpParser(tokens) ast = parser.parse() ``` Users of the {py:obj}`LP ` component therefore do not need to manage the lexer and parser separately.