How it was made
Every mapping was drawn before it was believed
The 14 pre-Unicode ways of writing Amharic are recorded here, 2167 mappings in all. This page is how those mappings were decided — and, which matters more, what they do not establish.
ማስረጃ
What the vendor’s own file doesn’t say
Every one of these old programs came with a file saying which byte picks which box in a grid of letters. What it doesn’t say is which letter that box is. The program never needed to know — its job was to draw a shape, and the shape had no name.
All of the work is in that gap. Crossing from a box to a letter can be done by eye, and an eye gives you a result that looks clean and is wrong. So it wasn’t done by eye.
Every mapping was drawn twice
Each byte was drawn in the vendor’s own font. The letter it claims to be was drawn in a Unicode font. The two pictures were compared. Where they matched, the mapping was published; where they didn’t, it was left out — not recorded as a near-miss, left out. That is why the converter tells you it couldn’t read a byte instead of offering its best idea: the cell that was left out did mean something, and nobody could establish what.
Every encoding, measured
For each encoding: the font it was checked against, the score that font reached, how many boxes came out of the vendor’s file, how many of those survived, and how many Amharic letters it reaches. The low scores at the bottom aren’t an accident — the small families carry only a handful of letters, and the rest of their boxes could not be established at all.
| Encoding | Checked against | Score | Boxes | Published | Letters |
|---|---|---|---|---|---|
| PowerGeez | Geez2.ttf | 0.996 | 329 | 298 | 280 |
| Nyala/NCI | ETSAMIB.ttf | 0.997 | 301 | 296 | 279 |
| VisualGeez | vg2title.ttf | 0.993 | 322 | 281 | 265 |
| VG2000 | VG2000T.ttf | 0.992 | 321 | 274 | 257 |
| NCI2000 | ETSAMIB.ttf | 0.996 | 301 | 272 | 255 |
| Alpas | Etsaba.ttf | 0.854 | 310 | 216 | 197 |
| Agafari (old) | Agafari_Rejim.ttf | 0.568 | 288 | 97 | 80 |
| Agafari (new) | Agafari_Rejim.ttf | 0.568 | 288 | 96 | 80 |
| Samawarfa | Addis98.ttf | 0.669 | 308 | 74 | 55 |
| Alex | ALXETHIO.ttf | 0.689 | 279 | 71 | 55 |
| Ethiopic PIC 1 | ETH-1.ttf | 0.555 | 185 | 53 | 36 |
| Fedel | ETH-1.ttf | 0.541 | 185 | 51 | 34 |
| Compose | ETH-2.ttf | 0.498 | 176 | 44 | 27 |
| Ethiopic PIC 2 | ETH-2.ttf | 0.498 | 176 | 44 | 27 |
Each encoding was tried against 39 fonts and the closest fit won. Often several score close together — and those turn out to be one encoding drawn at different weights, which is why a document asking for any of them reads the same way.
Where a person overruled the machine
Of the 2167 mappings, 2 rest on a person’s judgement rather than on the picture comparison. One is in PowerGeez, where the comparison picked the wrong vowel mark; the other is in Visual Geez, where the evidence is what somebody actually wrote. Both stories are on their own pages.
In both, the verdict the picture gave was kept. It sits in the record as it stood, with the decision that overruled it written beside it, so you can read what was overruled and why. Saying so doesn’t weaken the mapping — not saying so would.
Where the vendor’s own grid was wrong
The opposite happened too, 13 times: the vendor’s grid wrote one box twice, both claiming the same letter, and the rendering disagreed. There the rendering is what settled it. Both kinds of correction are recorded, and kept apart — one is a person overruling the machine, the other is the machine overruling the vendor.
What this does not establish
Nobody can repeat this work. The vendors’ driver files and fonts can’t be redistributed, so only the result is here. The strongest available claim is that these are the mappings that were derived — and every file is recorded with the fingerprint it had when it was copied, so a later edit to one of them shows.
42 Amharic letters are reached by no encoding at all. That is not an omission: the vendors’ grid has forty rows and these series are not among them, so they could never have been typed in these encodings in the first place. The converter reaches 284 of 326 — 87.1%.
Nor can the bytes alone tell you which encoding a document was written in. 14 pairs know the same bytes and read them as different letters, so both report reading the document completely. That is why the converter shows you what each reading makes of your file and leaves the choice with you.
The one thing not derived from the tables
Everything above is the tables agreeing with themselves. What breaks the circle is 2 real documents written by other people, with 6 phrases from inside them written down by somebody who reads Amharic. If the converter doesn’t produce those, something is wrong. Without them, the rest is a proof that checks itself.
Two records, and where they differ
There are two records of this work: a summary written onto each encoding itself, and a second file listing them all together. Both were written by one process, so their agreeing doesn’t establish that the work is right. Their disagreeing would establish that one of the two has changed since — which is the likeliest thing to go wrong here.
They agree on everything but one column. On the two Agafari encodings the second file counts one more discarded box than the first does — and those two are the only encodings with boxes the rendering could not draw at all. The difference is recorded rather than reconciled: nobody knows which count is right, and what is known is that it goes no further.