Maize under cultivation on smallholder plots outside Kaduna, Nigeria

AGBE

Àgbẹ̀ · farmer · an offline farm advisor

Bulletin No. 01 Africa Deep Tech Challenge 2026 Domain: Agriculture Gemma 3 1B · GGUF Q4_K_M

AGBE is a farming advisor you can ask questions, and it runs entirely on an ordinary laptop with the internet switched off. It is a small language model trained on West and Central African agronomy: crops, pests, livestock, soils, storage and getting produce to market. No cloud, no API key, no data cost.

Runs on
8 GB laptop
Memory used
1,039 MB
Speed
24.29 tok/s
Internet
None
Language
English

Try it

Three commands. Then pull out your network cable.

The weights are public and the runtime is llama.cpp, so nothing here is a demo you have to take on trust.

Terminal
curl -L -o agbe.gguf \
  https://huggingface.co/NEVODESIGN/agbe-1b/resolve/main/agbe-1b-q4_k_m.gguf

llama-cli -m agbe.gguf -t 4 -ngl 0 -c 2048 -st \
  -p "My maize has holes in the young leaves and wet sawdust in the whorl. What is this?"

814 MB download, once. After that it never touches the network again. Watch the demo · Model on HuggingFace · Source and model card

The problem

Advice exists. It just never reaches the field.

Nigeria has roughly one agricultural extension officer for every few thousand farming households. The knowledge that would raise a smallholder's yield is not secret and it is not new. It is written down in extension manuals, and it does not travel the last mile, because the last mile has no officer and often no signal.

The obvious answer is a farming chatbot. That breaks the moment you look at where farmers actually are. Rural coverage is patchy, mobile data is a real cost paid out of a thin margin, and a tool that needs the network is a tool that is absent on the morning the armyworm arrives.

Offline is not a feature we added. It is the shape of the whole thing: if it cannot answer with the network off, on hardware a farmer or a co-operative already owns, it does not count.
A smallholder plot intercropping maize with cassava
Plate 1 Maize intercropped with cassava. The questions AGBE is trained on are the ones this plot raises: spacing, weeding windows, what the yellowing on the lower leaves means, and whether it is worth spraying at all.

What it does, and what it refuses

Both halves matter.

It will

  • Identify a pest from what you can see in the field, and tell you how to confirm it before spending money
  • Give spacing, timing and rotation advice for the crops actually grown here
  • Talk you through poultry, goats, catfish, soils, drying and storage
  • Hold a conversation, so "what if I cannot afford that" gets a real answer
  • Say when it does not know, rather than inventing a pest or a product

It will not

  • Give you an agrochemical dose. Rates differ by product, and a confident wrong number is dangerous
  • Quote you a market price. Prices move weekly and an invented one costs you money
  • Advise on human health. Asked about a sick child it declines and points to a clinic
  • Pretend a virus has a cure. Cassava mosaic has none, and it says so

The challenge

What was actually asked for, and how it is marked.

The Africa Deep Tech Challenge 2026, run by the Africa Deep Tech Foundation, asks for a small language model that runs offline on a budget laptop and is genuinely useful in one domain. We entered Agriculture, one of seven tracks, against a field of about 1,700 participants.

Target machine8 GB RAM, integrated graphics, no GPU, Ubuntu 22.04
Must runFully offline, through llama.cpp, from a bare GGUF
Judged onSubmitted and hidden prompts evaluated by the organisers
Our trackAgriculture, English, African Use Case claimed

The hidden prompts are the interesting constraint. They exist to catch a model tuned to its own submitted examples, so there is no way to prepare for them except to be broadly competent. That single rule shaped the whole corpus.

Where the 100 marks sit Half is a human reading the answer. We could only engineer the other half. 50Accuracy 30Throughput 20Memory Judged by humans, on submitted and hidden prompts Relative to the fastest entry Linear, 7 GB budget −10 Thermal penalty, if the processor throttles or passes 85°C. Applied after everything else, so it is the cheapest ten points to lose.
Figure 1 The two shaded bars are the only parts a developer controls directly, and the second is scored against the fastest submission, not a fixed bar. Reading that before writing any code is what pointed us at a 1B model.

Why a 1B model, not the biggest that fits

The scoring formula tells you what to build, if you read it.

Scoring
Stotal = 0.50·Sacc + 0.30·Sperf + 0.20·Seff − Pthermal

Sperf = 100 × (TPS_act ÷ TPS_max)
Seff  = max(0, (7.0 − peak RAM GB) ÷ 7.0) × 100

Memory is paid for linearly, so every gigabyte is charged. Throughput is scored relative to the fastest submission, with 15 tok/s published only as a provisional reference, so surplus speed is not wasted the way a fixed cap would make it. So the instinct to pick the largest model that fits inside 8 GB is precisely backwards. We measured five candidates on the target profile rather than guessing.

Stacked bar chart comparing five candidate models at selection time, scored under the published formula. Qwen2.5 0.5B banks 48.6 of the 50 engineering points, Gemma 3 1B 34.8, and Qwen2.5 3B only 18.1, losing on throughput and memory at once.
Figure 2 Points available before any human accuracy judging. The 3B is the instructive row: it has the lowest measured throughput of the candidates and its 3.26 GB footprint collapses the memory term. It concedes 16.7 points to Gemma 3 1B before answering a single question, losing on both terms at once.

The corpus

Nothing in the training data is unattributable.

No dataset ships with this challenge, so the corpus is the real work. Scraping the web or having a large model write the answers would both have been faster, and both put claims into the data that nobody can trace. This domain is graded by agronomists who notice invented chemistry.

Every training pair is composed from a curated fact base of established extension practice. If a fact is not in that file, it cannot appear in the corpus.

  1. No invented numbers. No fabricated yields, no prices, and above all no agrochemical doses. Where a real answer needs a rate, it points at the product label and the local extension officer.
  2. Local grounding. Cassava mosaic, fall armyworm, striga, aflatoxin, Newcastle disease. Crops, varieties and seasons a farmer in the middle belt actually uses.
  3. Multi-turn. Judges hold a live conversation, so the model is trained on follow-ups: what if I cannot afford that, is it too late in the season, I tried that and it came back.
  4. Refusals are trained, not bolted on. Roughly one example in thirteen teaches the model where its competence ends.

Honest status

Where this actually is, as of today.

The engineering is measured: 814 MB on disk, 1,039 MB peak memory, 24.29 tokens per second on CPU with no GPU. The weights are published and the download script is verified end to end.

Throughput and memory are settled and neither depends on temperature: 24.29 tokens per second and 1,039 MB of the 7 GB budget, which gives 85.5 on S_eff. Against the provisional 15 tok/s reference that is 47.1 of the 50 engineering points, but S_perf is scored relative to the fastest submission, so our share of it depends on what everyone else ships.

One number depends on a rule we do not yet know. The official profiler runs a heavy 512-token prompt-processing pass on top of generation, and on this i7-10850H that reaches 100°C and throttles. We tested whether preparation helps: one run from 88°C and one from a genuinely cold 44°C peaked within a degree of each other and both throttled. If the penalty is taken from our telemetry it is 37.1; if the audit re-measures in its own sandbox, as the profiler's schema implies when it notes that cloud hosts usually report no thermal sensor at all, it is 47.1.

If the thermal penalty is judged…S_perfS_effP_thermalTotal
in the audit sandbox10085.5047.1
from our own telemetry10085.5−1037.1
Both figures are honest. Throughput and memory are identical either way: 24.29 tok/s and 1,039 MB of the 7 GB budget. Only the penalty differs. We quote 37.1, the figure we can prove on our own hardware.

The penalty is a property of this laptop rather than the model. An i7-10850H is a 45 W part in a thin chassis; under the profiler's 512-token prompt pass it throttles, while the same model in ordinary generation at four threads peaks at 83°C and does not. Pinning to four cores to match the target machine made it worse, not better: throughput fell below the provisional 15 tok/s reference and it throttled anyway. The judging FAQ describes a sandbox "resource-capped to match the Standard Laptop profile", and a properly cooled datacentre host would not throttle under the same load, so the sandbox figure is likely the higher one. We have asked the organisers which rule applies.

The final v13 build was evaluated on a 66-prompt behaviour battery and a 92-prompt hostile battery: 49/66 overall, 79/92 hostile, 56/62 attacks withstood, and zero safety leaks.

A model like this is only ever validated in a field, not in a benchmark.

Limits

Stated plainly, because a tool that hides its edges is worse than one that names them.

A 1B model is a knowledgeable extension pamphlet that can hold a conversation. It is not an agronomist, it is not a vet, and it is emphatically not a doctor. It knows West and Central African smallholder systems and will be thinner outside them. Where a real extension officer is available, they are the better answer, and AGBE is built to say so.

Questions

The things people actually ask, answered plainly.

Using it

Why Gemma 3 1B and not something larger?

Because the scoring formula punishes size and we read it before writing code. Memory is charged linearly against a 7 GB budget, so every gigabyte is paid for, and a 3B model concedes points before it answers anything. Throughput is scored relative to the fastest submission rather than against a fixed bar, which we initially misread as a cap; the technical notes record what that misreading cost us.

Why Q4_K_M quantisation?

It gives 814 MB on disk and about 1 GB in memory, which is the balance point between answer quality and the memory term in the score. It is the quantisation the model-selection curve was measured on, so the comparison between candidates is like for like, and it is what the published weights are.

Why not use retrieval instead of fine-tuning?

Retrieval would have been a reasonable design, and on a larger machine it might be the better one. We fine-tuned because the judged artefact is a single GGUF file run through llama.cpp, with no application layer around it, so a retrieval index would not be part of what gets scored. Given that constraint, the knowledge had to live in the weights.

Can I reproduce your results?

Yes, and that is the intent. The corpus generator, the training script and the Kaggle notebook that runs the whole pipeline end to end on a free T4 are all in the repository. The weights are public. Every number quoted on this site comes from a tool you can run yourself.

Scoring and measurement

Why do you quote two different scores?

Because one input to the formula depends on a rule we do not yet know. Throughput and memory are settled and identical either way, giving S_perf 100 and S_eff 85.5 at the provisional 15 tok/s reference; the final S_perf depends on the fastest submission. The only variable is whether a ten point thermal penalty attaches to us, which depends on whether the organisers measure temperature from a participant's own telemetry or inside their audit sandbox. We quote 37.1, the figure we can prove on our hardware, and show 47.1 as what the same model earns on a machine that is not thermally constrained.

Why does your laptop overheat?

It is a 45 W processor in a thin chassis. Under the profiler's 512-token prompt-processing pass it reaches 100°C and throttles, and it does so from a cold start with the case elevated and a fan running. In ordinary generation at four threads, which is what a farmer asking questions actually produces, the same model peaks at 83°C and takes no penalty. We also tried pinning to four cores to match the target machine and it made things worse: throughput fell below the provisional 15 tok/s reference and it throttled anyway.

Are the numbers on this site measured or estimated?

Measured, on the target profile, and the figures come from the official ADTC profiler rather than from our own tooling. Where the two disagreed we published the official one: our harness reported 0.88 GB of memory and the profiler reported 1,039 MB, so the site quotes the profiler. The raw submission.json is committed to the repository.

The wider picture

Who is this actually for?

Smallholder farmers and agricultural extension officers in West and Central Africa. Nigeria has roughly one extension officer for every few thousand farming households, and the districts with the fewest officers also tend to have the weakest connectivity. Those two facts together are the reason this runs offline rather than as an app.

Is this meant to replace extension officers?

No, and it would be a poor substitute. It is meant to be available on the days when no officer is, which for most smallholders is most days. The model is explicitly trained to defer to a human where the stakes are high, which is the opposite of trying to replace one.

What is next?

Three things. Fix the tail drift, which is the main remaining defect. Rebuild Nigerian Pidgin properly, with a corpus designed for it rather than a slice bolted on. And get it into a field, because a model like this is validated by an extension officer using it for a season, not by a benchmark.

What licence is it under?

The code and corpus are open in the repository. The base model is Gemma 3 1B, used under the Gemma Terms of Use, and inference runs on llama.cpp.