A Unicode/JSON layer around the Morpheus Ancient Greek morphological analyzer. Morpheus itself speaks beta code and prints Perseus-format analyses; this package makes it usable on arbitrary Greek texts:
- Unicode in/out — converts UTF-8 Greek to beta code (and back) with
beta-code. - Structured output — parses Morpheus's analyses into records with lemma, fine-grained POS, and morphological features; emits JSON.
- Ranking — re-orders candidate analyses by lemma frequency (and optional
context). Morpheus sorts candidates alphabetically by lemma, which puts the
wrong reading first surprisingly often (e.g.
θεοῦ→ the verb θεάομαι). - Tokenization — splits raw Greek text into tokens.
- Alignment — uses the engine's
-qdelimiter so misses stay in position.
Build the analyzer first (from the repository root):
scripts/build.sh # -> bin/cruncherThen install the toolkit:
python3 -m venv .venv
.venv/bin/pip install -e python/ pytest# raw UTF-8 text
.venv/bin/morpheus-toolkit analyze text.txt --json out.jsonl
# a columnar token file (MorphGNT / Apostolic Fathers); token is column 5
.venv/bin/morpheus-toolkit analyze data/morph/011-didache.txt --morphgntFor a whole corpus, scripts/analyze_corpus.py keeps the reference/gold columns,
writes JSONL, and prints a coverage summary:
scripts/analyze_corpus.py ../apostolic-fathers/data/morph/*.txt \
--ref-field 1 --token-field 5 --lemma-field 7 --lang-field 8 --lang-value grcThe Apostolic Fathers (Greek, 63,222 tokens) come out at 98.2% analyzed.
from morpheus_toolkit import Morpheus
morpheus = Morpheus() # uses ./bin/cruncher + bundled frequencies
for result in morpheus.analyze_text("ἐν ἀρχῇ ἦν ὁ λόγος"):
best = result.best
print(result.token, best.pos, best.lemma, best.feature_string())Morpheus parses Latin as well as Greek. Set language="latin" (CLI:
--language latin) and the tokens are passed through untouched (no beta-code
round-trip):
Morpheus(language="latin").analyze_text("femina amo dominus")Every Greek analysis carries a Strong's number (from the SBLGNT lemma mapping) and a softmax confidence; Latin analyses have neither:
best = morpheus.analyze_text("θεοῦ")[0].best
best.strongs # 2316
best.confidence # 0.32The Strong's table is MorphGNT-convention, so a few Morpheus spelling variants (γίγνομαι vs γίνομαι) resolve through their lemma variants when possible.
The bundled frequency table is derived from the MorphGNT SBLGNT (CC-BY-SA 4.0). Rebuild it with:
scripts/build_lemma_frequencies.py --morphgnt-dir /path/to/morphgntMeasure agreement against a gold corpus with scripts/agreement_report.py:
scripts/agreement_report.py --morphgnt-dir /path/to/morphgnt # SBLGNT
scripts/agreement_report.py ../apostolic-fathers/data/morph/*.txt \
--lemma-field 7 --pos-field 2 --source-field 9 --lang-field 8 --by-sourceMeasured effect of ranking (gold = the corpus lemma; "top1" = gold is the first candidate; "before" = Morpheus order, "freq" = after frequency ranking):
| Corpus | tokens | coverage | top-1 before | top-1 freq | POS |
|---|---|---|---|---|---|
| SBLGNT (MorphGNT) | 137,554 | 93.7% | 86.7% | 93.0% | 76.6% |
| Apostolic Fathers (all) | 63,222 | 90.9% | 84.5% | 89.9% | 74.6% |
| — of which MorphGNT-sourced | 52,094 | 96.4% | 89.7% | 95.7% | 75.9% |
Ranking adds ~6 points of top-1 lemma accuracy over Morpheus's native order,
out of domain. The lower numbers on the grc_proiel_lg subset reflect the noise
in that (machine-generated) reference more than Morpheus itself.
Two contextual approaches were built and measured but do not beat plain frequency ranking, so both are off by default:
- a POS-bigram Viterbi tagger (
tagger.py,--tagger): 92.9% vs 93.0%; - joint (lemma, POS) frequency (
--joint): 91.9% top-1 but 77.8% POS (vs 76.2%).
They need a learned emission model (P(lemma | tag)) and a proper tagger to pay off; the bundled hand-built model is not enough.
- POS granularity. Morpheus tags nouns, adjectives, pronouns, articles and
prepositions all as
N; the stem type recovers most of them, but noun vs adjective remains ambiguous. POS agreement is therefore lower than lemma agreement. - Contextual ranking is experimental. The hand-written context rules in
ranking.pyhurt accuracy and the POS-bigram Viterbi tagger intagger.pymerely matches frequency ranking; both are off by default. - Coverage. ~1.6% of NT tokens get no analysis, dominated by biblical proper
names (Δαυίδ, Φαρές, Ἀμιναδάβ, …) absent from
stemlib. Enableunknown_as_proper=True(CLI:--unknown-as-proper) to give unanalyzed tokens a synthetic proper-noun analysis (lemma = surface, POSN/NP, flaggedproposed=True). This brings the analysis rate to 100% on both the SBLGNT and the Apostolic Fathers. ~6% of tokens are analyzed but with a different lemma than the gold corpus.
cd python && ../.venv/bin/pytest -q