What we measured, what broke, and what thirteen model builds taught us about a 1B model's limits. Written as it happened, including the parts that did not work.
Read the scoring function before writing code
The challenge publishes its formula. Accuracy is judged by people; everything else is arithmetic.
Stotal = 0.50·Sacc + 0.30·Sperf + 0.20·Seff − Pthermal Sperf = 100 × (TPS_act ÷ TPS_max) Seff = max(0, (7.0 − peak RAM GB) ÷ 7.0) × 100 Pthermal = 10 if the CPU throttles or exceeds 85°C
Two things fall out, and one of them we initially read wrong. Memory is charged linearly, so every gigabyte is paid for. We first misread the provisional 15 tok/s reference as a cap, which would have made surplus speed worthless. The published formula is 100 × (TPSact ÷ TPSmax), so the denominator is the fastest submission and extra speed keeps paying. The instinct to run the largest model that fits in 8 GB is therefore backwards, and it is what most of a field of roughly 1,700 will do.
Measure five candidates, do not reason about them
We built a harness that holds the machine to the ADTC Standard Laptop profile: four threads, no GPU offload, memory capped to 7 GB, CPU-only llama.cpp.
The 3B concedes 16.7 points to Gemma 3 1B before answering a single question, which it would have to win back on accuracy alone. We chose Gemma 3 1B over the 0.5B believing it cost about one point. That was our own misreading of the provisional 15 tok/s reference: throughput is scored relative to the fastest submission, and under the published formula the choice cost 13.77 points, not one. We did not switch, because recovering that would need the 0.5B to lose fewer than 27.5 points of panel accuracy, and fitting agronomy into a 1B was the hard part of this project. The shipped figure elsewhere on this site is 1,039 MB, the official profiler's peak RSS for the final model: a different measurement of a different thing, and the one that counts for scoring.
A thermal result that runs backwards
Our first clean run lost the full ten-point penalty at 91°C. Sweeping thread counts produced a result that looked wrong:
| Threads | tok/s | Peak °C | Penalty | Points |
|---|---|---|---|---|
| 2 | 20.8 | 99.0 | −10 | 36.39 |
| 3 | 25.8 | 99.0 | −10 | 36.39 |
| 4 | 20.6 | 83.0 | 0 | 46.39 |
| 6 | 25.9 | 84.0 | 0 | 46.39 |
We initially misread the provisional 15 tok/s reference as a cap, which made surplus
throughput look free to trade for thermal headroom. In our own harness that appeared
to recover ten points. The cap was a misreading: the published formula is
100 × (TPSact ÷ TPSmax), so throughput is scored against the
fastest submission and surplus speed is never worthless. The thermal measurements below
stand; the trade we thought we were making does not.
The official profiler then took them back. It runs a 512-token prompt-processing pass as well as generation, a much heavier sustained load, and on this i7-10850H that hits 100°C and throttles from a 55°C cold start with the case elevated and a fan running. Pinning to four physical cores, since the Standard Laptop is a four-core machine, made it worse: throughput fell to 14.88 tok/s, below the provisional 15 tok/s reference, and it throttled anyway.
| If P_thermal is judged… | S_perf | S_eff | P_thermal | Total |
|---|---|---|---|---|
| in the audit sandbox | 100 | 85.5 | 0 | 47.1 |
| from participant telemetry | 100 | 85.5 | −10 | 37.1 |
The corpus, and why it is composed rather than scraped
No dataset ships with this challenge. Scraping the web or having a large model write the answers would both have been faster, and both put claims into the training data that nobody can trace. This domain is graded by agronomists who notice invented chemistry.
Every pair is composed from a curated fact base of established extension practice. If a fact is not in that file, it cannot appear in the corpus. No fabricated yields, no prices, and above all no agrochemical doses: where a real answer needs a rate, it points at the product label and the local extension officer.
Thirteen builds, and four different problems
Every build was judged by reading its answers, not by its loss curve. The loss curve looked healthy for every failure below.
| Build | Change | Result |
|---|---|---|
| v1 | rank 16, 3 epochs, 375 examples | Learned our answer scaffolding, not the agronomy. Told a parent to take a feverish child to an extension officer. |
| v2 | Removed repeated lead sentences; 568 examples | Stapled unrelated facts together, because the generator padded answers with other topics' facts. |
| v3 | Stopped cross-topic mixing; coherent fact base | Refusal and timing correct. Pest wrong ("pod borers"); Pidgin answered in English. |
| v4 | 42 fall-armyworm mentions, Pidgin to 6.7% | Invented "fall army weevil". More examples did not fix a confusion. |
| v5 | rank 32, 5 epochs | Facts finally correct in English and Pidgin. But coherence broke: word salad, and it invented a pesticide called "dorabacite". |
| v6 | rank 32, back to 3 epochs | English facts and refusal both correct and stable. Tail degeneration persists, and Pidgin named a food ("amala") as a pest. Pidgin claim withdrawn here. |
| v7–v8 | Corrective exemplars from a 66-prompt battery | Closed a jailbreak that had produced a paediatric paracetamol dose. Then answered "when should I plant maize" with oil palm spacing. |
| v9 | Sentence cap, planting calendars, 40 adversarial exemplars | Safest build made, 94% of 62 attacks withstood. But it lost facts: blossom end rot became "bacterial wilt". |
| v10 | 4 epochs, to recover those facts | Recovered some, and began inventing vocabulary: "mortjacket" for coccidiosis. The v5 failure again. Reverted. |
| v11 | Symptom-first diagnosis, rare facts protected from the cap | Diagnosis 7/12 → 10/12. The corpus had only ever asked about named diseases, never about symptoms. |
| v12 | Contrast exemplars for confusable livestock pairs | Gained one livestock prompt, gave back safety leaks. Rejected. |
| v13 | Bidirectional contrast, armyworm vs stem borer | Shipped. 49/66, zero safety leaks, and it names fall armyworm on our own submitted prompt where v11 and v12 both said stem borer. |
The pattern only became visible at v5, and it is the most useful thing this project produced:
Coherence is a third axis, and it degrades with over-training. v5 recalled facts and then filled the remaining tokens with fluent nonsense once it exhausted what it had learned. Adding examples cannot fix that, and neither can adding capacity. v10 proved the point twice by inventing "mortjacket" for coccidiosis five builds later.
Sentence diversity is a fourth axis, and we were not measuring it for six more builds. An audit after v8 found 1,020 conversations built from only 900 unique sentences: 5.6 reuses each, and one sentence appearing 40 times. At three epochs the model saw each one about seventeen times and memorised it as a unit, then emitted those units by topic rather than by question. That is how "roughly 143 palms per hectare", an oil palm figure, ended up inside an answer about maize. Growing the corpus from 642 to 1,020 conversations made the model worse, because the count grew and the diversity did not.
The last of these is the subtlest. Contrast exemplars fix confusions: naming the alternative and rejecting it worked where volume had failed. But a contrast teaches a direction, not a boundary. Five examples reading "that is stem borer, not armyworm" and none of the reverse left v12 answering the textbook fall armyworm description with "that is stem borer", on our own submitted test prompt. An unbalanced contrast relocates a bias rather than removing it, and the corpus builder now audits every such pair for its opposite.
A model that invents a pesticide name is more dangerous than one that admits it does not know. That is why refusal was trained first and guarded hardest.
Knowing when to stop, and what to withdraw
By the sixth build, English diagnosis and refusal behaviour were reliable. Nigerian Pidgin is not. It was correct in v5 and in v6 it named amala, a food, as a maize pest. Adding more Pidgin examples is exactly what we tried between v3 and v4 for pest naming, and it produced the invented "fall army weevil" rather than a fix.
So the Pidgin claim was withdrawn from language_scope rather than shipped and hoped for. The African use-case claim is unchanged, because it never rested on language: it rests on the domain being cassava mosaic, striga, aflatoxin, Newcastle disease and harmattan planting windows.
A capability that works in one build and invents food names in the next is not a capability. Claiming it would have been the easiest thing on this page to get away with, and the fastest way to lose a judge who speaks the language.
The environment, which fought back
Recorded because they cost real time and every one is reproducible:
peft 0.19.1 alongside torchao 0.10.0, and that peft raises on any torchao below 0.16 from inside its LoRA dispatcher.pip install -U pulled versions that no longer matched the preinstalled torchvision, and disturbed TensorFlow's protobuf. Both break import transformers outright. The fix was to stop upgrading and remove what is unused.trl broke inside its own chunked cross-entropy path on a PEFT-wrapped causal LM. We dropped it for plain transformers.Trainer, which also made the label masking explicit and verifiable.<image_soft_token> sits at id 262144 with vocab_size 262144, so llama.cpp's converter fails its vocabulary assertion at the very end, after writing every tensor.llama-cli hangs without -st. It generates the answer and then waits on a terminal that never comes, which cost a 900-second timeout per prompt before we found --single-turn.What we would tell the next team
Read the scoring function first; it told us to build a 1B when instinct said 3B. Measure on the target profile rather than reasoning about it, because the thermal result ran backwards from intuition. Judge a model by reading its answers, since every failure here had a healthy loss curve. And when style transfers but facts do not, the problem is capacity, not data.