@keboola/sql-splitter
v0.2.1
Published
Splits a SQL script into statements using the CodeMirror SQL grammar, so quotes, dollar-quotes and comments follow the dialect rather than a hand-written scan.
Downloads
107
Readme
@keboola/sql-splitter
Splits a SQL script into the fragments the SQL editor runs and stores.
import { splitStatements } from '@keboola/sql-splitter';
splitStatements("SELECT 'a;b';\nSELECT 2;", 'snowflake');
// ["SELECT 'a;b';", "\nSELECT 2;"]Why it is not a regex
Statement boundaries depend on which constructs a dialect treats as opaque. A regex that encodes them needs nested quantifiers over an alternation, which backtracks catastrophically as soon as a quote or a comment is left open — and the editor re-splits a document that is still being typed.
Boundaries here come from the CodeMirror SQL grammar (@codemirror/lang-sql), the same
tokenizer the editor already runs. Quotes, dollar-quotes, comment syntax and identifier
quoting follow the dialect definition instead of a hand-written scan, so '', "",
backslash escapes and nested block comments come for free.
Guarantees
- Total. Fragments always re-join to the input, including for unterminated literals
and comments. A script the splitter cannot make sense of is still a script the user can
run, and the backend gives a better syntax error than we can. A fragment is usually one
statement, but a script can also yield a bare
;or a trailing comment. - Linear. One pass over the input and one over its tokens. No input makes it backtrack, and the work per byte does not climb with the size of the script — there is a test for that shape, not for a wall-clock threshold.
Dialects
| | Snowflake | BigQuery |
| --------- | ----------------- | ----------------- |
| "…" | quoted identifier | string |
| `…` | not special | quoted identifier |
| $$…$$ | string | not special |
| # | not a comment | line comment |
| // | line comment | not a comment |
Tests
splitStatements.test.ts is the readable spec. behaviour.test.ts is the safety net: a
characterisation matrix of 1,440 cases per dialect — every lexical construct alone, next to
every separator, and paired with every other construct — snapshotted whole, so a change
anywhere shows up as a reviewable diff rather than a single failing assertion.
