Most machine translation engines treat Tetun as a low-resource language and hallucinate freely — fluent-sounding output with wrong words in it. This translator is built to minimise silent guessing: every sentence is checked against real Tetun evidence before you see it, and the parts that are already known don't involve AI at all. It can still make mistakes, so review anything important before you rely on it.
First, the translation memory
The translator keeps a memory of human-verified English–Tetun sentence pairs. If your exact sentence is already in it, you get the stored human translation back verbatim — instantly, no AI involved. If a nearly identical sentence is known, the AI is instructed to minimally edit that trusted translation rather than draft from scratch — the same way a professional translator works from translation memory. Single words with one unambiguous dictionary answer are served straight from the glossary.
For everything else: draft, verify, repair
- Draft. Claude drafts natural Tetun — phrased the way a Timorese writer would put it, not a word-for-word calque of the English — respecting your register (formal, neutral, casual) and domain (justice, health, government, tourism).
- Verify the terminology. Key terms are checked against a 28,000+ row glossary built from the Dili Institute of Technology (DIT) dictionary, DIT Justice Sector, DIT Health & Medical, DIT Tourism, DIT Wordfinder and IT glossaries, and the INL Matadalan Ortográfiku ba Tetun-Prasa (2003).
- Verify against real usage. The sentence is compared with 90,000+ parallel English–Tetun sentences from Timor-Leste government publications, DIT course materials, tetun.org, and Timorese civil-society writing. Where real human phrasing exists, it wins over anything invented.
- Lint. The draft is checked against 79 grammar and spelling rules — Portuguese loan endings (-saun, -u), aspect markers, negation, apostrophes, common English-calque patterns — with DIT or INL orthography enforced.
- Repair. If the linter still finds an error in the final text, an automatic repair pass fixes exactly the flagged spans — nothing else is touched. You will occasionally see a translation correct itself a moment after it appears; that is the repair pass working.
And it learns from Timor-Leste
Translations that keep being requested and pass every check are reviewed and promoted into the translation memory, where they serve future users instantly. The thumbs-down button — and the “know a better translation?” box behind it — feed a review queue, so corrections from fluent speakers make the translator better for everyone. The memory also grows from newly digitised bilingual sources: government text, DIT publications, and civil-society archives.
The linguistic rules are drawn from the Peace Corps Tetun Language Course (3rd ed., Catharina Williams-van Klinken, 2015), the DIT-TLPDP Tetun for the Justice Sector textbook (2015), and the INL orthography standard (Decree 1/2004). Machine translation still makes mistakes — for anything that will be signed, filed, or published, have a qualified human translator review the output.