🌿 fiche-extract v1.1.0 · MIT
Extract text and embedded media from an XLSX technical sheet — one pass, source file never modified.
What it does
- Text + images together:
openpyxlandsoffice --convert-to csvonly give you cell text — real-world spec sheets often carry critical info in embedded photos/diagrams. fiche-extract gets both. - One CSV per sheet via LibreOffice headless; every
xl/media/*entry extracted and inventoried (bytes, type). - Safe by construction: the source file is opened read-only; everything lands in a separate output folder.
- Explicit JSON report: sheets, line counts, media inventory, warnings (macros detected, unexpected formats, missing soffice…).
- Honest exit codes:
0content extracted ·1failure/no content ·2bad input (missing, not .xlsx, zero bytes).
Install (pip, v1.1+)
pip install git+https://mandrilly.com/git/fiche-extract.git
Installs a fiche-extract command (zero Python dependencies — uses your system soffice):
fiche-extract spec.xlsx # → ./fiche_extract_spec/{csv/, media/, report.json}
fiche-extract spec.xlsx -o out_dir
Quick start (from source)
python3 fiche_extract.py spec.xlsx # → /tmp/fiche_extract_spec/{csv/, media/, report.json}
python3 fiche_extract.py spec.xlsx -o out_dir
Source
git clone https://mandrilly.com/git/fiche-extract.git
Requires Python 3.8+ and LibreOffice (soffice). Test fixtures included.