When this translation was shared publicly, the most common objection was immediate and reasonable: a language model has read every famous English Odyssey (Lattimore, Fagles, Wilson, and the rest are in its training data), so its "translation" must really be a re-synthesis of theirs. One commenter called the opening "a crappy mishmash of Lattimore and Fagles"; another wrote that it "clearly combines a few translations line by line."
Part of that objection is untestable, and we concede it up front: prior translations live in the model's weights regardless of what it was prompted with, just as they live in the memory of any human translator working today. No output analysis can establish what a model, or a person, was shaped by. What can be measured is reuse: whether the text of this translation reproduces the wording of its predecessors, at what lengths, and how that compares with how much human translators of the same poem reproduce each other. This page reports those measurements, states exactly what they do and do not establish, and takes the accusation's strongest form seriously, including a word-by-word look at the very lines the accusation was made about.
Ten texts: this translation (~124,000 words) and nine English Odysseys spanning 293 years. The panel: Pope (1725), Cowper (1791), Butcher & Lang (1879), Palmer (1891), Butler (1900), Murray (1919, the Loeb prose facing the very Greek text this project translated), Lattimore (1967), Fagles (1996), and Green (2018). Each is trimmed to translation body only: prefaces, introductions, and footnotes removed, so that quoted or editorial matter cannot masquerade as translation (this matters; see "What copying looks like" below). Everything is lowercased, punctuation, possessives, and diacritics stripped, and proper names normalized across traditions (Ulysses→Odysseus, Lattimore's Telemachos→Telemachus, Green's Athēnē→Athena), so spelling conventions can neither hide nor simulate overlap.
For every pair of texts, two measurements: the fraction of one text's n-word sequences (n = 4…8) appearing anywhere in the other, taking the larger of the two directions; and every maximal verbatim shared passage of 12 or more words, with no cap on length. All 36 human-vs-human pairings are computed with the identical procedure and shown below, the smallest grouped for space. The controls are the whole point, since any two literal translations of the same Greek converge substantially without either copying the other. The Lattimore, Green, and Palmer texts come from scans; residual OCR noise systematically depresses comparisons involving those texts.
Public-domain texts come from Project Gutenberg and the Perseus/Scaife library; the script downloads them and prints a hash of each cleaned corpus. Lattimore, Fagles, and Green, still in copyright, were measured from privately held copies; only statistics are published, never their text. A public clone of the analysis script therefore reproduces the public-domain rows exactly and the copyrighted rows given your own copies.
| this translation vs | 5-gram overlap | 8-gram | runs ≥12 words | runs ≥16 | longest run |
|---|---|---|---|---|---|
| Murray (1919, literal prose) | 8.9% | 2.3% | 170 | 27 | 29 |
| Lattimore (1967, literal line-for-line verse) | 6.3% | 1.1% | 64 | 11 | 24 |
| Green (2018, literal line-matched verse) | 5.3% | 0.8% | 49 | 3 | 20 |
| Butcher & Lang (1879, literal prose) | 5.3% | 1.0% | 52 | 7 | 22 |
| Palmer (1891, literal prose) | 4.0% | 0.7% | 32 | 4 | 19 |
| Fagles (1996, free verse) | 1.7% | 0.2% | 5 | 0 | 14 |
| Butler (1900, free prose) | 1.6% | 0.1% | 0 | 0 | <12 |
| Cowper (1791, blank verse) | 0.2% | 0.0% | 0 | 0 | <12 |
| Pope (1725, couplets) | 0.02% | 0.0% | 0 | 0 | <12 |
And the complete human-vs-human control matrix, all 36 pairings of the nine prior translations, same procedure, sorted:
| human pair | 5-gram | human pair | 5-gram |
|---|---|---|---|
| Murray · Butcher & Lang | 15.2% | Butler · Green | 1.2% |
| Murray · Green | 5.0% | Lattimore · Fagles | 1.1% |
| Murray · Palmer | 5.0% | Butler · Palmer | 1.1% |
| Murray · Lattimore | 3.5% | Fagles · Green | 1.0% |
| Lattimore · Green | 3.2% | Palmer · Fagles | 0.8% |
| Butcher & Lang · Palmer | 2.8% | Murray · Fagles | 0.8% |
| Butcher & Lang · Lattimore | 2.3% | Butler · Fagles | 0.6% |
| Butcher & Lang · Green | 2.2% | Butcher & Lang · Fagles | 0.5% |
| Palmer · Green | 1.8% | Cowper · Palmer | 0.3% |
| Murray · Butler | 1.8% | Murray · Cowper | 0.3% |
| Palmer · Lattimore | 1.7% | Butcher & Lang · Cowper | 0.3% |
| Butler · Lattimore | 1.5% | remaining 5 Cowper pairs | ≤0.2% |
| Butler · Butcher & Lang | 1.4% | remaining 7 pairs (all involve Pope) | ≤0.04% |
Run counts for the closest human pairs: Murray · Butcher & Lang share 492 runs of ≥12 words (105 of ≥16, longest 32); Murray · Palmer 52 (10, longest 20); Murray · Green 42 (4, longest 21); Murray · Lattimore 21 (2, longest 21); Green · Lattimore 17 (0, longest 15); Lattimore · Fagles 3 (0, longest 13).
Read plainly, the tables say four things:
Overlap tracks literalness of method, in humans and machine alike. Among humans, the literal translations occupy the entire top of the matrix while Cowper and Pope, translating the identical poem, share 0.1%. This translation's affinities follow the same gradient: highest with the most literal predecessors, lowest with the poets. That profile is consistent with independent literal translation, and not with a pastiche of the famous poetic versions.
The matched-method human control behaves like the machine. Green's 2018 translation (literal, modern register, line-matched to the Greek, the closest existing human analogue to this project's method) shows the second-highest human affinity in the whole matrix (5.0% with Murray), and nobody supposes Green copied the Loeb. This translation's relationship to Green is numerically very similar to Green's own relationship to Murray: 5.3% vs 5.0% at 5-grams, 49 vs 42 long runs, longest 20 vs 21. The model sits with the line-faithful literalists where a new member of that method would sit. Green also supplies the table's sharpest test between the two explanations for overlap, because fame and method point in opposite directions across his row and Fagles's. Training-data prevalence cannot be measured directly; but to the extent that quotation, anthologizing, and course adoption proxy it, Fagles is among the most-reproduced English Odysseys ever published and Green's 2018 version among the least-quoted modern ones. Memorization predicts affinity should follow that prevalence; method-convergence predicts it should follow literalness. The observed result: three times the affinity for the little-quoted methodological sibling (5.3%) as for the ubiquitous poetic one (1.7%). Century-old literalists nobody reads (Butcher & Lang, Palmer) also outscore Fagles. Within this panel, affinity follows method, not fame. This is an inference over the table, not a measurement, and it shares the table's limits.
Two rows remain elevated, and we say so. Against Murray (8.9%) and Lattimore (6.3%), this translation runs roughly twice Green's own affinity to the same texts (5.0%, 3.2%), higher than every human pairing except Murray · Butcher & Lang, a single towering pair at 15.2% that is itself no certificate of independence (Murray postdates Butcher & Lang by forty years in the same archaizing tradition; the median human pairing is about 0.5%). Two explanations fit the residue and these numbers cannot decompose them: this translation is stricter than even Green (same line count, same word-order discipline, and Murray's own Greek text as source, all of which push convergence up), and the model may carry some real gravity toward the literal translations in its training data. What the numbers do bound is the form any such influence took: the run profile stays human-shaped throughout (11 runs of ≥16 words with Lattimore, against the Murray · Palmer human pair's 10), and verbatim reuse is capped by the next section.
The Fagles signal is modest but, in fairness, not absent. This translation's 1.7% overlap with Fagles exceeds every human-vs-Fagles pairing in the panel (the highest are Lattimore's 1.1% and Green's 1.0%). That is worth stating because it is the kind of detail a defense would prefer to omit. But it is a fraction of the Murray and Lattimore affinities, it comprises five shared runs of 12+ words and none of 16+, and the longest match (14 words) is a Homeric formula that Fagles also shares verbatim with Murray. The data do not support Fagles as a meaningful textual donor; they cannot rule out minor influence.
The "stitching" version of the accusation (that the text was assembled from pieces of prior translations) was tested against the union of all nine at once: an n-gram counts as matched if it appears in any of them. Result: 79.5% of this translation's 5-grams and 95.1% of its 8-grams appear in none of the nine; at 4 words, 65.8% appear nowhere. The longest passage shared with any predecessor, anywhere in 12,107 lines, is 29 words (with Murray, the facing translation of its own source text).
What do those figures actually exclude? To calibrate the test rather than assert about it, we built the thing being alleged, at several scales: synthetic splices cycling verbatim chunks of Murray, Butcher & Lang, and Lattimore, with chunk sizes from 4 to 12 words.
| text | 4-grams found in union | 5-grams found | runs ≥12 words vs Murray |
|---|---|---|---|
| splice of exact 4-word chunks | 26.6% | 0.9% | 0 |
| splice of exact 5-word chunks | 41.3% | 20.7% | 0 |
| splice of exact 6-word chunks | 51.1% | 33.9% | 0 |
| splice of exact 8-word chunks | 63.3% | 50.5% | 0 |
| splice of exact 12-word chunks | 75.6% | 67.0% | 3,490 |
| this translation | 34.2% | 20.5% | 170 |
The calibration cuts both ways, and both directions belong on this page. Against long-fragment assembly it is decisive: splices at 8 and 12 words score two and a half to three times this translation's figure on the 5-gram measure (1.9–2.2× at 4-grams), so the results strongly disfavor substantial assembly from verbatim passages of roughly eight to twelve words or longer. But a text built entirely of exact five-word fragments scores 20.7% on the 5-gram test, essentially this translation's 20.5%, because most sliding windows cross fragment boundaries. N-gram union statistics cannot exclude short-fragment mosaic composition at the 4–6 word scale, and we do not claim they can.
One measured observation bears on that remaining hypothesis without settling it. In these cyclic splices, sub-12-word chunks produce no shared runs of 12+ words, though a differently built mosaic could if adjacent fragments came from the same source. This translation shares 170 such runs with Murray and 49 with Green, a long-tailed profile that resembles the human controls (Green shows 42 against Murray) more than it resembles the synthetic splices. That is suggestive of convergence, not diagnostic of it: a differently constructed mosaic or a mixed text could produce another profile, and nothing in these statistics rules that out. Finer-grained dependence is a question for reading, not string-counting, which is why the next section reads the actual lines the accusation cited.
The accusation was made about the proem specifically: that its first line "combines" Lattimore's man of many ways and Fagles's man of twists and turns into man of many turnings. N-gram statistics can't adjudicate a three-word phrase; the Greek can.
ἄνδρα μοι ἔννεπε, μοῦσα, πολύτροπον, ὃς μάλα πολλὰ / πλάγχθη, ἐπεὶ Τροίης ἱερὸν πτολίεθρον ἔπερσεν
man [acc.] · to-me · tell · Muse · much-turning [acc.] · who · very much · was-driven-to-wander · after · of-Troy · holy · citadel · he-sacked
This translation renders: "Tell me the man, Muse — the man of many turnings, who was driven / wandering far, once he had sacked Troy's holy citadel." What follows reads the line both ways: where it argues from the Greek, and where it genuinely resembles the accused sources.
"Tell me the man" keeps Homer's bare accusative: ἄνδρα is the direct object of ἔννεπε. Every alleged source inserts an "of": Murray "Tell me of the man," Lattimore "Tell me, Muse, of the man," Fagles "Sing to me of the man." (Green feels the same pull the Greek exerts; he fronts the noun, "The man, Muse — tell me about that resourceful man," but resolves it with "about.") The one construction in the line that is uniquely this translation's is the one closest to the Greek. Against that: the line then repeats "the man" where the Greek has a single ἄνδρα, the same doubling Fagles uses (and Green, in his own way) to bridge the long gap between noun and epithet. A defensible solution to a real problem, but a solution Fagles arrived at first.
"Man of many turnings": πολύτροπος is πολύ ("many") + τρόπος ("turn"). "Turnings" is the root-literal rendering, closer to the Greek than Lattimore's "ways," Fagles's "twists and turns," Murray's "devices," or Green's "resourceful," and identical to none of them. The "man of many ___" frame is shared with Murray and Lattimore, half a century apart. Greek attaches a compound adjective to "man," English has no adjective "much-turned," and the genitive frame is the standard literalist's escape (though not the only one: Butcher & Lang found "so ready at need," Palmer "the adventurous man").
"Driven wandering far, once he had sacked": πλάγχθη is the passive of a verb meaning "drive off course," which is why "driven" appears in Lattimore and Fagles too; "driven wandering" spends two English words unpacking the one Greek verb, a defensible but not inevitable choice. "Once he had" for ἐπεί matches Fagles's construction ("once he had plundered") where Murray and Lattimore write "after"; either is ordinary English for the clause, and this translation's word order follows the Greek's, but the resemblance is there to see.
"Troy's holy citadel" shares its frame with Lattimore's "Troy's sacred citadel" and swaps the adjective; ἱερόν can be either. Divergence inside a shared construction is consistent with independent translation of the same words; it is not, by itself, evidence of independence, and we don't claim otherwise. And a concession critics counted fairly: μάλα ("very") goes untranslated in line 2, the kind of small-word sacrifice to meter and idiom that human translators make on every page.
The honest summary of the proem: these are the most Lattimore-and-Fagles-adjacent lines anyone has identified in the poem: the most translated, most quoted, most memorized hexameters in Greek, where the gravitational pull of famous renderings is at its strongest. They contain real echoes alongside choices independently warranted by the Greek. What a single conspicuous line cannot do, though, is sustain a verdict about a 124,000-word text; that is what the whole-book measurements above are for, and they bound long-form reuse tightly.
The corpus supplied its own demonstration. Butler's Project Gutenberg file contains a 253-word verbatim match with Butcher & Lang, because Butler's preface quotes their rendering of the proem at length in order to mock it. Once prefaces and notes are trimmed away and only translation bodies compared, the Butler/Butcher & Lang figure collapses to three shared runs, longest 14 words. That is the difference between quotation and convergence: real reproduction announces itself in hundreds of consecutive words on the first string search. Nothing remotely like it exists between this translation and any predecessor.
Where translations do converge verbatim, they converge on the same lines: Homer's formulas. The longest passage this translation shares with Fagles and the longest it shares with Green are the same passage: the recurring clothing-promise ("…a two-edged sword, and sandals for his feet, and send him wherever his heart and spirit bid…"), a formula the poem itself repeats five times (14.516 → 21.339) and which this translation, following its rule that repeated Greek formulas recur verbatim in English, renders identically at each return, as Murray, Fagles, and Green do too. The places that look most "copied" are the places Homer copied himself.
A cross-poem aside: against Fagles's Iliad (his voice, a different poem) this translation shares 0.30% of 5-grams (longest run: 8 words, all stock formulas), of the same order as the 0.20% that Murray's 1919 Odyssey shares with it. No formal test of style is claimed here; the point is only that no gross Fagles house-style signal appears.
What these measurements cannot do: establish what shaped the model (conceded at the top); detect dependence carried in short fragments (the calibration above shows the four-to-six-word scale is outside this test's reach), in syntax, or in interpretive choices, the debts every translator, human or machine, owes their predecessors; or say anything about literary quality, copyright, or the ethics of machine translation, which are real questions outside this page's scope. Wilson (2017) and Fitzgerald (1961) are not yet in the panel; Wilson's row is expected soon, and predictions for it are preregistered in the repository, commit-timestamped before any measurement, with falsification conditions stated in advance. Line-aligned comparison and syntactic or semantic similarity measures would probe finer dependence than n-grams can; they remain open work.
Provenance: this page was prepared with the same class of model that made the translation, so don't take our framing on trust: the method is fully specified above, the script downloads and hashes the corpora, and a public clone reruns the entire table (copyrighted texts supplied by path). Before publication the analysis was audited by a model from a different lab (the same reviewer lineage used in this project's two-agent review passes), which surfaced a run-length cap bug, preface contamination in two corpora, and framing that outran the metrics. The corrections are incorporated above and itemized in the repository's commit history; Green was added to the panel afterward, closing the matched-method gap the audit identified.