@bidilens/core
v0.3.0
Published
Unicode direction analysis, isolation, streaming state, and bidi security primitives.
Maintainers
Readme
@bidilens/core
Framework-independent Unicode direction analysis, inline-isolation planning, streaming state, and bidi security auditing. Runtime analysis is offline and uses generated Unicode 17.0.0 bidi-class and natural-letter data rather than the host runtime's potentially older Unicode property tables.
npm install @bidilens/coreimport {
analyzeText,
createBidiStream,
planInlineIsolation,
sanitizeBidiControls,
scanBidiSecurity
} from '@bidilens/core';
const source = 'React یک کتابخانه جاوااسکریپت بسیار محبوب است.';
const analysis = analyzeText(source);
// analysis.direction === 'rtl'
// analysis.text === source
const isolations = planInlineIsolation(source, 'rtl');
// React is an LTR identifier isolation.
const security = scanBidiSecurity(source, { mode: 'strict' });
const sanitized = sanitizeBidiControls(source, {
removeGroups: ['embedding-override', 'deprecated']
});
// Directional marks and complete isolate/PDI pairs are preserved.
// With removeGroups, risk filtering applies to the whole atomic family.
const stream = createBidiStream();
stream.push('old response');
stream.reset(source); // clear and analyze replacement text atomicallyneedsBidiIntervention(text, context) exposes the shared non-interference
gate. LTR-only text in an LTR context returns false; RTL content, bidi
formatting controls, an inherited RTL context, or { intervention: 'always' }
returns true. Inline planning likewise returns no unnecessary ranges for
ordinary LTR-only text.
The default newline paragraph separator is recognized incrementally. A custom
paragraphSeparator regular expression is evaluated once by finish() so
future-sensitive lookarounds, anchors, and extendable matches remain invariant
across arbitrary source chunking. Until then, custom-separated input is exposed
as one unresolved open paragraph. Set paragraphBoundary: 'markdown' to
recognize blank lines incrementally while retaining a single soft line break
inside the current paragraph. If both options are supplied, the explicit
Markdown boundary policy takes precedence in core and framework rendering.
The default content-majority policy excludes technical tokens before
counting natural-language evidence. Raw-text analysis recognizes closed
multiline backtick/tilde fences embedded in surrounding prose; standalone and
incomplete fences retain their literal evidence. Parser-aware Markdown
integrations remain the preferred path for structured documents. Use
first-strong or strict-uax9 only when compatibility with first-strong host
behavior is required.
Built-in recognition covers conservative, unambiguous tool and product names. Add private or domain-specific single-token identifiers without changing global state:
analyzeText('internalplatform خوب است.', {
technicalIdentifiers: ['InternalPlatform']
});Matching is case-insensitive. Custom values must start with an ASCII letter and
may then contain ASCII letters, digits, _, ., or -. firstStrong reports
the first strong character after configured technical-token exclusion;
rawFirstStrong reports the literal first Unicode bidi-strong character.
The default stream direction remains provisional and revisable until
finish(). Choose sticky-majority only when UI stability is more important
than correcting a misleading prefix before completion.
Run the packaged example after building with
pnpm --filter @bidilens/core example.
