๐ฟ transcript-fixer v1.0.0 ยท MIT
Deterministically fix proper names that speech-to-text mangles โ before the transcript reaches any downstream consumer.
The problem
ASR output (e.g. gpt-4o-transcribe) renders recurring proper names badly โ and prompt-level hints do not fully solve it, because the error is acoustic, not lexical. The fix is a deterministic rewriting pass: a curated, sourced table of mis-transcribed โ correct forms applied with word boundaries and idempotence guarantees.
What it does
- Sourced fix table โ each entry backed by a real observed ASR error; edit it to match your own recurring mistakes.
- Idempotent โ
fix(fix(x)) == fix(x); safe in multi-pass pipelines. - Protected forms โ already-correct names are never rewritten (no false positives on close words).
- Word boundaries, case-insensitive, longest multi-word expressions first.
- CLI + Python API:
--json,--list,--selftest, stdin/file input. - Zero dependencies, Python โฅ 3.9, single module.
Install (pip)
pip install transcript-fixer @ git+https://mandrilly.com/git/transcript-fixer.git
Usage
transcript-fixer "I flew to Nu York with Micheal" # โ I flew to New York with Michael [2 fix(es)] echo "Open Ai and Git Hub" | transcript-fixer - transcript-fixer --selftest # 7/7 OK transcript-fixer --json report.txt # machine-readable, with change details
Python API
from fix_names import fix_names, fix_report
fixed, changes = fix_names("Send the file to Micheal")
# ("Send the file to Michael", [{"from": "Micheal", "to": "Michael", "pos": 17}])
Tests
transcript-fixer --selftest # 7/7 โ OK (fixes, idempotence, protection)
Source
git clone https://mandrilly.com/git/transcript-fixer.git