Writing
A herbarium for texts
Most of what we call a knowledge system is really a laundering machine. Sources go in; confident, unattributed claims come out. Somewhere in the middle the provenance is quietly discarded, disagreements are resolved by whoever held the pen last, and a guess becomes a fact by the simple act of being restated without a hedge. I have spent enough time with the texts of ancient south-west Asia to distrust any system that makes knowledge look tidier than it is. Tidiness, in this domain, is usually a symptom of something lost.
The model I keep returning to is the herbarium. A herbarium does not tell you what a plant means. It pins the specimen to a sheet, verbatim, and labels it with where it was collected, when, and by whom. The specimen is the ground truth; it is never edited to match a later theory. Observations about the specimen — its classification, its relationships, its supposed range — are recorded separately, and each observation points back to the pinned original. You can always walk from the interpretation to the thing interpreted. That directionality is the whole discipline.
Apply this to text and a few rules fall out that are more demanding than they first appear.
Preserve the original verbatim, with its provenance attached, and never silently overwrite it. A correction is a new layer, not an erasure of the old one. Record derived observations separately from the source and make every derivation cite the original it rests on. Never invent an attribution, and never invent a provenance link to paper over a gap — a fabricated citation is worse than an admitted absence, because it survives scrutiny that the truth would not.
Unknown has to be a first-class value
The hardest discipline is admitting ignorance in the data model itself. Most systems have no way to say “I don’t know” — they have present values and missing values, and missing quietly reads as false, or gets filled by whatever the algorithm finds plausible. But in this material, “the referent of this pronoun is genuinely unrecovered” is not a defect to be smoothed over. It is the finding. A system that cannot represent uncertainty will manufacture certainty, and manufactured certainty is indistinguishable from a lie once it has propagated a few hops downstream.
Then there is disagreement, which is where flattening does the most damage. A passage attributed to one hand by one scholar and to another by a second is not a conflict to be adjudicated by the database. Both attributions are real facts about the state of the question. An honest system holds them side by side, each with its source, and refuses to collapse them into a single winner. The disagreement is information. Resolving it prematurely destroys that information and disguises a live scholarly question as a settled one.
This is the thinking behind Codexarium, but I would defend it with no reference to any project at all. The claim is narrower and more stubborn than any tool: a system that cannot show you where a statement came from, cannot say when it does not know, and cannot hold two people disagreeing is not preserving knowledge. It is laundering it, and doing so all the more convincingly for looking so clean.
#scholarship#open knowledge