01The industry is drowning in data and starving for knowledge
Somewhere in Honduras, a coffee farm has been surveyed three times in two seasons. An exporter's agronomist walked the plot boundaries and logged them. A certifier came later and recorded the same boundaries again for a different form. A buyer's sustainability team asked for the altitude, the cultivars, the shade trees and the number of harvest workers — all of which had already been written down twice.
Three visits. Three datasets. Zero compounding.
Each of those records sits in a separate system, in a separate format, owned by a separate company, answering a separate question. The farmer cannot see most of it. The exporter cannot use the certifier's version. The roaster who eventually buys the coffee gets a one-line origin description and writes a story around it. And when that farmer sells to a different buyer next season, the whole exercise starts again from zero.
This is the actual condition of coffee knowledge in 2026. Not scarcity — the industry collects an enormous amount. The problem is that almost none of it accumulates. Every actor pays to gather facts that someone else already gathered, stores them where no one else can reach them, and derives from them a fraction of the value they contain.
The European Union's Deforestation Regulation is about to make this much more expensive. It applies from 30 December 2026 for large and medium operators, with small and micro enterprises given until 30 June 2027, and it requires tracing coffee back to the plot where it grew. Every operator in the chain is now building a private version of the same map. The duplication is not a side effect. It is the design.
We think there is a better structure available, and it is not another platform. It is a Coffee Knowledge Vault: a shared, public, machine-readable base of coffee knowledge that many parties deposit into and everyone can draw on. This article is an argument for why the industry needs one, an account of what each participant gets out of it, and an invitation to help build the first version.
02Why the value is trapped
The 2026 Coffee Barometer, published in June, is the most useful reality check available right now. Its argument is that the recent price boom exposed coffee's structural problems rather than solving them.
The dependence on smallholders is not in dispute. Roughly 12.5 million coffee-growing households supply the world. About 95% farm plots smaller than five hectares. Farms under two hectares are around 80% of all coffee farms and produce roughly 60% of global supply. Yet in eight of the ten largest producing countries, the average coffee-farming household still earns below a living income.
Prices did not fix that. Arabica futures hit an all-time high near $4.40 per pound in February 2025, and the Barometer records the ICO composite indicator setting a record monthly average of 354.52 cents. By March 2026 that indicator had fallen to 273.70 cents, more than 22% off the peak; as of mid-August 2026 arabica trades around $3.17 per pound. Farmgate prices rise more slowly than retail prices and fall more sharply when the market turns. Producers absorb the volatility; everyone downstream is better positioned to hold a margin.
Sjoerd Panhuysen, one of the Barometer's authors, put a warning on the table that anyone building technology for this sector should sit with:
"Technical solutions are almost always easier than political or economic ones. It is far simpler to improve traceability, train farmers or design a certification scheme than it is to change how value, risk and power are distributed across the chain."
He is right, and we are not claiming a knowledge vault fixes prices. What it does address is a different and genuinely tractable problem: an enormous amount of value in this industry is currently destroyed by fragmentation rather than captured by anyone. The roaster who cannot find a lot matching a flavour profile, the farmer who cannot learn what worked on the next ridge, the cooperative that cannot prove its quality improvement, the researcher who cannot get a dataset, the importer paying to re-document a farm already documented — none of those losses show up on anyone's balance sheet as a loss. They just quietly reduce what the whole sector is worth.
That is the value a shared knowledge base creates. Not by redistributing an existing pie, but by making knowledge compound instead of evaporate.
03What a Coffee Knowledge Vault is
A vault, in the sense we mean it, is a knowledge commons: a structured, verifiable, openly governed body of knowledge that many parties contribute to and all can query.
The idea is not new outside coffee. Genetics has GenBank. Geography has OpenStreetMap. Chemistry, astronomy, and structural biology all run on shared reference bases that no single company owns and that made entire fields dramatically more productive. In each case the same thing happened: once the shared layer existed, competition moved up the stack. Nobody competes on whether a gene sequence is correct. They compete on what they do with it.
Coffee has no such layer. It has certifications, private traceability systems, trade statistics, academic papers and an enormous amount of undocumented expertise sitting in the heads of farmers, millers and roasters. None of it is connected.
A Coffee Knowledge Vault would hold, in one connected structure:
- Farms and farmers — location, area, altitude, coordinates, cultivars, ownership, the cooperative or group
- Lots — each harvest batch, linked to the plot it came from
- Practice — processing method, fermentation parameters, drying, who milled it
- Quality — cupping scores, physical and chemical measurements, defects
- Sensory vocabulary — flavour descriptors from published wheels, including regional ones, mapped to each other
- Movement — collection, storage, shipping, roasting, the chain from plot to cup
- Evidence — the sensor readings, lab results and documents that back the claims above
- Reference knowledge — the agronomic and coffee-science layer that gives all of the above meaning
The first deposits are the simplest ones: a farmer registering their own farm. Name, farm name, where it is, how big, how high, what is planted, who processes the cherry. Every other layer in the vault attaches to that record. It is the foundation not because it is the most sophisticated data but because everything else is meaningless without it.
04Why this is possible now and was not before
Shared reference bases used to take a decade to negotiate. A standards committee argued about field definitions, published a specification, and then discovered that nobody in the field could be bothered to fill in the forms.
Three things changed.
Ontology engineering matured.
We now know how to build a formal model of a domain that is rigorous enough for machines and extensible without breaking what came before. Adding a new stakeholder or a new concept no longer means renegotiating the whole schema.
Graph databases made relationships first-class.
In a property graph, the connection between a farm's altitude and a lot's cupping score is not something you reconstruct in a query. It is stored, traversable, and cheap to follow.
LLMs removed the data-entry barrier.
This is the decisive one. The reason old knowledge bases failed was that contributing to them was miserable work, and the evidence on that is brutal: only about 10% of small farmers use agricultural apps even when they are free. The reasons are mundane and instructive — shared phones, changing SIM cards, information delivered once in language that is hard to parse. A farmer will not fill in a forty-field form, and no amount of good intention changes that. A farmer will type or say a sentence about what they did today, in their own language, and a capable model will turn that sentence into structured knowledge that lands correctly in the graph.
That finding shaped our architecture more than any other. The entry point is a messaging app people already use daily, not an application to download. Input is free-form text or voice, not a form. When the system is unsure it asks one short question rather than silently guessing. And the original message is preserved before anything is interpreted.
The aim is to keep typing to an absolute minimum. Photograph the drying beds in the afternoon. Record a voice note saying how long the ferment ran. Take a short clip while the pile is being turned. Or let the meter at the collection point report its own readings. Every one of those routes reaches the vault — the system does the work of turning them into structured knowledge, not the person.
Nobody should have to become a data-entry clerk to get their own knowledge on the record.
Structure used to be cheap and contribution expensive. That has inverted. Contribution is now easy; getting the structure right is the hard and valuable part.
The picture above is the whole idea in one frame. Every box is a thing that exists — a person, a place, a batch of coffee, a measurement, a term from a controlled vocabulary. Every line has a name and a direction. A conventional database would store the farm's altitude in one column and the cupping score in another table, and the connection between them would live only in whatever query someone happened to write. Here the connection is the data.
That is what makes a question like "farms at a similar altitude and soil to mine — which processing method gave the highest cupping scores?" answerable at all. To answer it a system walks from a farm to its conditions, out to comparable farms, from those farms to their lots, from each lot to its processing method and its score, and back with a ranking. That is a multi-hop relational query. Keyword search cannot do it. Vector similarity cannot do it reliably either, because retrieving passages that sound related is a different operation from following a verified chain of relationships.
We call this approach mission-driven ontology design, and it has one governing rule: nothing enters the schema without a competency question behind it — a real question a real participant needs answered. If no one is asking it, we do not model it. That rule is what keeps a knowledge base from becoming an elegant, expensive, unused taxonomy, which is how most of them have died.
There is a practical note for anyone worried about cost. In a property graph, query cost comes from the volume of instance data and the depth of traversal, not from how rich the schema is. A broad, well-designed structure with a modest amount of data is fast and cheap to run. You can be generous with structure and disciplined with data at the same time.
05What each participant deposits, and what each one withdraws
A commons only works if the exchange is real for everyone in it. Here is our account of it, stated plainly enough to be argued with.
Farmers
deposit their farm record and what they actually do each season — processing choices, fermentation times, drying, harvest notes. They withdraw comparison against farms with conditions like theirs, which is the single most valuable thing a smallholder cannot get today; a documented quality history; and a provenance profile that lets them ask a higher price with evidence behind it.
Cooperatives
deposit aggregation and quality-control data from the collection point. They withdraw a quality dashboard across members, the ability to demonstrate improvement over seasons, and a stronger position in negotiation because their claims are backed rather than asserted.
Processors and mills
deposit processing parameters and what happened to each batch. They withdraw the evidence base that connects process to outcome — the thing that turns a mill's craft from an unprovable claim into a demonstrable competence.
Roasters
deposit cupping results and sourcing outcomes. They withdraw verified lot profiles for the coffee they buy, the ability to search for lots by flavour profile or farm condition rather than by broker relationship, and a story for their customers that is true and checkable.
Cafés and consumers
deposit attention and, sometimes, sensory feedback. They withdraw traceability that is real rather than decorative, and the ability to follow a specific bag back to a specific plot and person.
Exporters and importers
deposit logistics and compliance events. They withdraw the plot-level data EUDR requires, collected once by the farmer with consent instead of re-collected by every buyer in the chain — which converts a recurring compliance cost into a one-time contribution.
Researchers and institutions
deposit agronomic science, reference vocabularies and published findings. They withdraw a real-world, structured, consented dataset of smallholder practice and outcomes that does not currently exist in any open form anywhere.
And everyone withdraws the compounding. The tenth roaster's cupping data makes the first farmer's benchmark better. The hundredth farm record makes every altitude comparison sharper. This is the property that makes a commons worth building rather than a product worth buying: the value of your deposit grows because of deposits other people make.
06Trust is what makes deposits from strangers usable
A shared vault has a problem a private database does not. If contributions come from many parties with different interests, why should anyone believe what is in it?
The answer is not to trust contributors more. It is to record how we know each thing.
A statement from a person is a claim, stored with its provenance: who said it, when, extracted from which original message, by which model, at what confidence. The original message is stored before extraction, always. If the model misread it, or the extraction improves later, the contributor's actual words are still there and can be reprocessed. Nothing anyone contributed is lost to a parsing error.
A sensor reading — near-infrared moisture at a collection point, satellite deforestation monitoring — is evidence. Its interpretation is a separate observation claim.
When claim and evidence agree, the claim is promoted from asserted to verified. When they conflict, it is flagged as contradicted, and nothing is deleted. A flag does not mean someone lied. Instruments drift and processes vary. The system marks an inconsistency for a human to look at; it does not pass judgment.
This is the mechanism that makes an open vault usable commercially. A roaster does not have to trust a farmer they have never met. They can see that a claim about honey processing was corroborated by a moisture reading, and decide accordingly. Verification, not credulity, is what lets strangers transact on shared knowledge.
We also make a deliberate choice about instruments. Not on every plot. Sensors belong at bottlenecks — the cooperative's collection point, the satellite pass — where one device validates hundreds of claims. Per-farm sensor deployment has never made economic sense for smallholder agriculture and still does not.
07Vocabulary is a contribution, not a constraint
One design decision looks small and turns out to matter a great deal for a commons.
The standard flavour wheel is built on reference fruits that make sense to a North American or European palate: blueberry, blackcurrant, cranberry. Ask a farmer at 1,200 metres in Chiang Rai to describe her coffee in those terms and you have asked her to describe her own crop in a language borrowed from a place she has never been.
So the vault supports pluggable, versioned flavour schemes with curated mappings between them. A Thai farmer describes her coffee in lychee, longan, mango, tamarind and roselle; a buyer who needs the international vocabulary reads it in mapped equivalents. Both are correct. Neither loses precision.
Framed as a commons, this stops being a local accommodation and becomes a contribution. A Thai or Asian flavour scheme is a piece of sensory knowledge the coffee world does not currently have in structured form. So is an Ethiopian one, and a Colombian one. Every origin that describes its coffee in its own terms and maps those terms into the shared structure makes the whole vault more expressive.
The general principle runs well beyond flavour: a knowledge system that forces people to describe their world in someone else's vocabulary collects worse knowledge. That is not a cultural nicety. It is a data quality argument.
It helps to picture the vault geographically. Any mapping tool can already show where a farm, a washing station, a warehouse and a roastery sit. What it cannot show is that this sack in that warehouse came from those three plots, was processed on a particular day, and carries a moisture reading consistent with what the miller recorded. Switch the knowledge layer on and the pins stop being pins. They become a chain anyone can walk in either direction — forward from a plot to the café that poured it, or backward from a bag on a shelf to the farm and the person.
Every new participant is a new node type attaching to a Lot that is already in the graph. That is the payoff of designing the ontology properly: the second stakeholder is far cheaper to add than the first, and the tenth is nearly free.
08How a commons stays a commons
Knowledge commons fail in predictable ways. They get captured by one large contributor, or they fragment into incompatible forks, or they fill with unverifiable junk, or the stewarding organisation quietly turns them into a product. Any serious proposal has to say how it intends not to do that.
Our commitments are four.
The structure is open.
The ontology, the schema and the reference vocabularies are published. Anyone can implement against them, extend them, or argue with them. A commons whose structure is proprietary is not a commons.
Contributors keep rights over their own records.
A farmer's farm record is theirs: portable, exportable, shareable with the buyers they choose, revocable from the ones they do not. This is where the ownership question belongs — as a governance rule that keeps participation voluntary and keeps the vault honest, rather than as the headline argument.
Provenance is mandatory, everywhere.
Every fact in the vault carries its source and its verification status. That is what prevents both capture and junk: a large contributor cannot overwrite the record, they can only add claims that stand or fall on their evidence.
Stewardship is separated from ownership.
Neo Gens funds and maintains the work. Neo Gens does not own the knowledge, and the design is deliberately unattractive to acquire, because the valuable part is published.
None of this is charity, and we should be plain about it. We think a functioning knowledge commons makes the whole specialty coffee sector more productive, and organisations that know how to build and use structured knowledge will do well in that world. That is the business we are in. But the vault itself has to be a public good, because a vault that one company owns is just a database with better marketing, and nobody should deposit into it.
09Proof: the first deposits
Global ambitions deserve local evidence, so here is an honest account of where the work stands.
We chose Thailand for the first deposits for a reason that surprised us. Thailand grows roughly 16,000 tonnes of coffee a year and drinks around 95,500 — it imported about 80,000 tonnes in 2025 and exported some $9.2 million worth of green and roasted coffee in 2024, ranking 87th in the world. Thai smallholders sell domestically, into a market that wants six times what the country produces. EUDR compliance is not their problem.
That made Thailand a better test, not a worse one. If a knowledge vault only attracts contributions when a regulation forces them, it is a compliance product. In Thailand the only reason anyone would contribute is that the vault gives back more than it takes. That is a much harder bar and a much more honest one.
What exists today: a formal domain ontology built on published academic work on Thai coffee knowledge and extended with supply-chain and data-integrity modules; a Neo4j schema with Lot as the central linking entity; two flavour vocabularies — the international wheel and a Thai/Asian wheel — connected by curated mappings; an eight-stage capture pipeline design from a messaging app to a validated graph write; the claim-and-evidence provenance layer described above; and a competency-question register in which six priority questions are implemented as queries and pass their evaluation tests against a demonstration dataset. The first tangible output, a one-page lot profile for a roaster, has been generated end to end. The ontology is built federation-ready, so a farm registered in Colombia, Ethiopia or Honduras joins the same structure as one in Chiang Rai.
What does not exist yet: contributors. The pilot cooperative, the field interviews that will test our competency questions against what growers actually ask, and the cupping sessions that will validate the Asian flavour wheel are the next phase, not a completed one. Every question in the register is still marked draft, precisely because it came from research rather than from a farmer's mouth. We are not promoting any of them until we have sat with people who grow coffee.
We would rather publish that distinction clearly than round it up. The sector has enough pilots described in the past tense.
10An invitation
A vault with one contributor is a database. What we are asking for is the first cohort.
- Cooperatives and producer organisations willing to make the first deposits in a second origin, and to tell us what the structure gets wrong about how coffee actually works where they are
- Roasters who will contribute cupping data and use farmer-sourced profiles in real sourcing decisions — and say plainly when the profiles are not good enough to use
- Mills and exporters who see the case for depositing processing and compliance data once instead of documenting the same farms every season
- Researchers, institutes and origin bodies with agronomic knowledge or a regional flavour vocabulary that belongs in a shared structure rather than in a PDF
- Development organisations and funders interested in knowledge infrastructure as a complement to the harder work on prices and contracts, not a substitute for it
- Ontology and knowledge-graph practitioners willing to read the schema and argue with it
The coffee industry already pays for all of this knowledge. It pays several times over, in duplicated surveys, unprovable quality claims, sourcing decisions made on thin information, and expertise that retires with the person who held it. The only thing missing is a shared place to put it and a structure that makes it compound.
That is a thing a group of people can decide to build. We would like to help build it, and we would like company.
Coffee KM is a public-goods project stewarded by Neo Gens, a Modern Knowledge Management practice. The ontology, schema and vocabularies are published openly. To contribute, get in touch.
Sources
- Inside the 2026 Coffee Barometer, Part 1: Prices Swing, Structures Don't — Daily Coffee News, July 2026
- 2026 Coffee Barometer — Ethos Agriculture et al., June 2026
- EUDR is Accelerating Farm Modernization, But Threats to Smallholders Remain — Daily Coffee News / Mongabay, July 2026
- EU Deforestation Regulation 2026 Update: New Deadlines for Companies — PSQR
- Is coffee ready for the EUDR? — FoodNavigator, April 2026
- "Error 404 farmers not found" — Berta Ortiz, Alliance of Bioversity International and CIAT
- Lertkrai, Kaewboonma & Wathanti (2023). Developing an Ontology for Thai Coffee Knowledge
- Phaekhiao, Khankasikam & Nuntawong (2025). A Comparison of the Effectiveness of Ontology-Driven Information System Development Tools, ICIC Express Letters 19(2), 221–230
- Coffee trade data — Thailand — OEC