pi-encoding-fs
v0.2.0
Published
Encoding-aware read/write/edit/grep for Pi — transparent GB18030/GBK/GB2312 support
Maintainers
Readme
pi-encoding-fs
Encoding-aware read / write / edit / grep for Pi.
Transparently handles GB18030 / GBK / GB2312 files while preserving Pi's native TUI
(diff view, syntax highlighting, line numbers).
How it works
The extension registers tools with the same names as Pi's built-ins, overriding them.
read/write/edit reuse Pi's own tool skeletons via createXxxToolDefinition, injecting
encoding conversion only at the byte-IO layer, so Pi keeps computing and rendering diffs in
UTF-8. grep is self-implemented so it can search Chinese text inside GB-encoded files.
Encoding is decided per file: the extension walks up from the file's directory to the
nearest .encoding-converter.json. If none is found, it passes through as plain UTF-8
(identical to not having the extension installed).
Configuration
Place .encoding-converter.json in any directory. It applies to files at or below it,
unless a deeper config overrides it.
{
"sourceEncoding": "GB18030",
"confidenceThreshold": 0.8,
"overrides": [
{ "pattern": "docs/**", "sourceEncoding": "UTF-8" },
{ "pattern": "legacy/**", "sourceEncoding": "GBK" }
]
}sourceEncoding: default encoding when detection is uncertain; also used for new files.confidenceThreshold: minimum chardet confidence (0-1) to trust auto-detection.overrides: glob rules relative to the config file's directory; most specific wins.
Requirements
- A Pi agent (peer dependency;
>=0.80.0). - Encoding detection (for
read/write/edit) uses Python 3 +chardetwhen available. If Python is missing, the extension silently falls back tosourceEncodingfrom config. grepuses ripgrep (rg) for fast, encoding-aware search. Ifrgis not found onPATH,greptransparently falls back to Pi's built-in grep (which cannot search inside GB-encoded files).
Install
From npm:
pi install npm:pi-encoding-fsFrom git:
pi install git:github.com/15wtyuan/pi-encoding-fsNotes / limitations
- Only
read/write/edit/grepare overridden.bash,find, andlsare not touched (e.g. acatinsidebashwon't decode GB files). grepshells out toripgrepwith--encodingper encoding group (derived from the nearest.encoding-converter.json), so searching Chinese inside GB18030 files is both correct and fast. It excludes.git/.svn/.hg/node_modulesby default. Honoring.gitignoreand aligning match limits with Pi's built-in grep are planned.
License
MIT
