← AGBE

Build notes

What we measured, what broke, and what thirteen model builds taught us about a 1B model's limits. Written as it happened, including the parts that did not work.

Africa Deep Tech Challenge 2026 Domain: Agriculture Gemma 3 1B · LoRA r32 · GGUF Q4_K_M

01

Read the scoring function before writing code

The challenge publishes its formula. Accuracy is judged by people; everything else is arithmetic.

Scoring
Stotal = 0.50·Sacc + 0.30·Sperf + 0.20·Seff − Pthermal

Sperf = 100 × (TPS_act ÷ TPS_max)
Seff  = max(0, (7.0 − peak RAM GB) ÷ 7.0) × 100
Pthermal = 10 if the CPU throttles or exceeds 85°C

Two things fall out, and one of them we initially read wrong. Memory is charged linearly, so every gigabyte is paid for. We first misread the provisional 15 tok/s reference as a cap, which would have made surplus speed worthless. The published formula is 100 × (TPSact ÷ TPSmax), so the denominator is the fastest submission and extra speed keeps paying. The instinct to run the largest model that fits in 8 GB is therefore backwards, and it is what most of a field of roughly 1,700 will do.

02

Measure five candidates, do not reason about them

We built a harness that holds the machine to the ADTC Standard Laptop profile: four threads, no GPU offload, memory capped to 7 GB, CPU-only llama.cpp.

Stacked bar chart of five models by engineering points.
Figure 1 Points available before any human accuracy judging, scored under the published formula. Qwen2.5 0.5B banks 48.6 of 50 and Gemma 3 1B 34.8. Qwen2.5 3B banks 18.1: it has the lowest measured throughput of the candidates and its 3.26 GB footprint collapses the memory term, so it loses on both terms at once.

The 3B concedes 16.7 points to Gemma 3 1B before answering a single question, which it would have to win back on accuracy alone. We chose Gemma 3 1B over the 0.5B believing it cost about one point. That was our own misreading of the provisional 15 tok/s reference: throughput is scored relative to the fastest submission, and under the published formula the choice cost 13.77 points, not one. We did not switch, because recovering that would need the 0.5B to lose fewer than 27.5 points of panel accuracy, and fitting agronomy into a 1B was the hard part of this project. The shipped figure elsewhere on this site is 1,039 MB, the official profiler's peak RSS for the final model: a different measurement of a different thing, and the one that counts for scoring.

03

A thermal result that runs backwards

Our first clean run lost the full ten-point penalty at 91°C. Sweeping thread counts produced a result that looked wrong:

Threadstok/sPeak °CPenaltyPoints
220.899.0−1036.39
325.899.0−1036.39
420.683.0046.39
625.984.0046.39
Fewer threads ran hotter. With only two or three cores loaded the CPU boosts toward its single-core turbo ceiling and per-core temperature spikes. At four or more the all-core power limit caps clocks, spreading the same work cooler. More parallelism, less heat, more throughput.

We initially misread the provisional 15 tok/s reference as a cap, which made surplus throughput look free to trade for thermal headroom. In our own harness that appeared to recover ten points. The cap was a misreading: the published formula is 100 × (TPSact ÷ TPSmax), so throughput is scored against the fastest submission and surplus speed is never worthless. The thermal measurements below stand; the trade we thought we were making does not.

The official profiler then took them back. It runs a 512-token prompt-processing pass as well as generation, a much heavier sustained load, and on this i7-10850H that hits 100°C and throttles from a 55°C cold start with the case elevated and a fan running. Pinning to four physical cores, since the Standard Laptop is a four-core machine, made it worse: throughput fell to 14.88 tok/s, below the provisional 15 tok/s reference, and it throttled anyway.

If P_thermal is judged…S_perfS_effP_thermalTotal
in the audit sandbox10085.5047.1
from participant telemetry10085.5−1037.1
Throughput and memory are identical in both rows and neither depends on temperature: throughput is 24.29 tok/s, and 1,039 MB of a 7 GB budget gives S_eff 85.5. Only the penalty moves. The judging FAQ describes a sandbox resource-capped to the Standard Laptop profile, and a cooled datacentre host does not throttle the way a thin laptop does, so the sandbox row is the likelier one. We quote the lower figure because it is the one we can prove: three profiler runs, from 88°C, 44°C and 54°C, all peaked at 99 to 100°C and all throttled. Starting cold does not help on this chassis, so the penalty is a property of the hardware rather than of the preparation.

04

The corpus, and why it is composed rather than scraped

No dataset ships with this challenge. Scraping the web or having a large model write the answers would both have been faster, and both put claims into the training data that nobody can trace. This domain is graded by agronomists who notice invented chemistry.

Every pair is composed from a curated fact base of established extension practice. If a fact is not in that file, it cannot appear in the corpus. No fabricated yields, no prices, and above all no agrochemical doses: where a real answer needs a rate, it points at the product label and the local extension officer.

05

Thirteen builds, and four different problems

Every build was judged by reading its answers, not by its loss curve. The loss curve looked healthy for every failure below.

BuildChangeResult
v1rank 16, 3 epochs, 375 examplesLearned our answer scaffolding, not the agronomy. Told a parent to take a feverish child to an extension officer.
v2Removed repeated lead sentences; 568 examplesStapled unrelated facts together, because the generator padded answers with other topics' facts.
v3Stopped cross-topic mixing; coherent fact baseRefusal and timing correct. Pest wrong ("pod borers"); Pidgin answered in English.
v442 fall-armyworm mentions, Pidgin to 6.7%Invented "fall army weevil". More examples did not fix a confusion.
v5rank 32, 5 epochsFacts finally correct in English and Pidgin. But coherence broke: word salad, and it invented a pesticide called "dorabacite".
v6rank 32, back to 3 epochsEnglish facts and refusal both correct and stable. Tail degeneration persists, and Pidgin named a food ("amala") as a pest. Pidgin claim withdrawn here.
v7–v8Corrective exemplars from a 66-prompt batteryClosed a jailbreak that had produced a paediatric paracetamol dose. Then answered "when should I plant maize" with oil palm spacing.
v9Sentence cap, planting calendars, 40 adversarial exemplarsSafest build made, 94% of 62 attacks withstood. But it lost facts: blossom end rot became "bacterial wilt".
v104 epochs, to recover those factsRecovered some, and began inventing vocabulary: "mortjacket" for coccidiosis. The v5 failure again. Reverted.
v11Symptom-first diagnosis, rare facts protected from the capDiagnosis 7/12 → 10/12. The corpus had only ever asked about named diseases, never about symptoms.
v12Contrast exemplars for confusable livestock pairsGained one livestock prompt, gave back safety leaks. Rejected.
v13Bidirectional contrast, armyworm vs stem borerShipped. 49/66, zero safety leaks, and it names fall armyworm on our own submitted prompt where v11 and v12 both said stem borer.
Four corpus iterations preceded any change to the training configuration, and rank was the missing variable the whole time. Seven builds later the same shape recurred: we were counting conversations when what the model actually sees is sentences.

The pattern only became visible at v5, and it is the most useful thing this project produced:

Behaviours

  • Refusing out-of-scope questions, hedging on doses, admitting no cure exists
  • Generalise from few examples
  • Fixed at rank 16, stable ever since

Facts

  • Which pest, which symptom, which spacing
  • Need model capacity, not more examples
  • Only fixed when LoRA rank doubled to 32

Coherence is a third axis, and it degrades with over-training. v5 recalled facts and then filled the remaining tokens with fluent nonsense once it exhausted what it had learned. Adding examples cannot fix that, and neither can adding capacity. v10 proved the point twice by inventing "mortjacket" for coccidiosis five builds later.

Sentence diversity is a fourth axis, and we were not measuring it for six more builds. An audit after v8 found 1,020 conversations built from only 900 unique sentences: 5.6 reuses each, and one sentence appearing 40 times. At three epochs the model saw each one about seventeen times and memorised it as a unit, then emitted those units by topic rather than by question. That is how "roughly 143 palms per hectare", an oil palm figure, ended up inside an answer about maize. Growing the corpus from 642 to 1,020 conversations made the model worse, because the count grew and the diversity did not.

The last of these is the subtlest. Contrast exemplars fix confusions: naming the alternative and rejecting it worked where volume had failed. But a contrast teaches a direction, not a boundary. Five examples reading "that is stem borer, not armyworm" and none of the reverse left v12 answering the textbook fall armyworm description with "that is stem borer", on our own submitted test prompt. An unbalanced contrast relocates a bias rather than removing it, and the corpus builder now audits every such pair for its opposite.

A model that invents a pesticide name is more dangerous than one that admits it does not know. That is why refusal was trained first and guarded hardest.

05b

Knowing when to stop, and what to withdraw

By the sixth build, English diagnosis and refusal behaviour were reliable. Nigerian Pidgin is not. It was correct in v5 and in v6 it named amala, a food, as a maize pest. Adding more Pidgin examples is exactly what we tried between v3 and v4 for pest naming, and it produced the invented "fall army weevil" rather than a fix.

So the Pidgin claim was withdrawn from language_scope rather than shipped and hoped for. The African use-case claim is unchanged, because it never rested on language: it rests on the domain being cassava mosaic, striga, aflatoxin, Newcastle disease and harmattan planting windows.

A capability that works in one build and invents food names in the next is not a capability. Claiming it would have been the easiest thing on this page to get away with, and the fastest way to lose a judge who speaks the language.

06

The environment, which fought back

Recorded because they cost real time and every one is reproducible:

  1. Kaggle's base image contradicts itself. It ships peft 0.19.1 alongside torchao 0.10.0, and that peft raises on any torchao below 0.16 from inside its LoRA dispatcher.
  2. Upgrading made it worse. pip install -U pulled versions that no longer matched the preinstalled torchvision, and disturbed TensorFlow's protobuf. Both break import transformers outright. The fix was to stop upgrading and remove what is unused.
  3. trl broke inside its own chunked cross-entropy path on a PEFT-wrapped causal LM. We dropped it for plain transformers.Trainer, which also made the label masking explicit and verifiable.
  4. Gemma 3 ships a token the text-only 1B cannot use. <image_soft_token> sits at id 262144 with vocab_size 262144, so llama.cpp's converter fails its vocabulary assertion at the very end, after writing every tensor.
  5. llama-cli hangs without -st. It generates the answer and then waits on a terminal that never comes, which cost a 900-second timeout per prompt before we found --single-turn.
  6. Truncated downloads pass a header check. A partial GGUF keeps a valid magic number at byte 0, so only the length reveals it. Our download script now verifies length and resumes.
  7. A CPU-only torch wheel arrives uninvited. llama.cpp's convert requirements install one, and once it lands the GPU build is gone for the rest of the container. A later re-run of the training cell then trains on CPU: no crash, no hang, just a 30x slowdown indistinguishable from a deadlock. The trainer now aborts rather than warns.
  8. And that fix broke the export, which took seven attempts to trace. Filtering torch out meant writing the requirements elsewhere, which broke a relative include inside the file, so pip aborted silently. That install had been quietly doing a second job: pinning transformers down from 5.0 to 4.57. Without the pin, conversion dies on a vocabulary assertion, because transformers 5.0's Gemma tokenizer injects a token at the one id the text-only 1B cannot hold. Five fixes were written for a tokenizer that was never broken.
  9. A checksum caught what a size check would have passed. Resuming a download onto a file already at full length appends nothing and exits happily, leaving old content at exactly the right byte count. The published weights are now verified by downloading all 814,261,088 bytes and hashing them, not by trusting a response header.

07

What we would tell the next team

Read the scoring function first; it told us to build a 1B when instinct said 3B. Measure on the target profile rather than reasoning about it, because the thermal result ran backwards from intuition. Judge a model by reading its answers, since every failure here had a healthy loss curve. And when style transfers but facts do not, the problem is capacity, not data.