Goodfire releases Silico tool for debugging models from inside
On April 30, MIT Technology Review reported that the startup Goodfire released Silico, a tool that lets engineers inspect individual neurons and pathways in a trained model and adjust connected parameters to strengthen or weaken behaviors, including during training. The article says Silico can trace pathways upstream and downstream of a neuron, works only with models whose internals are accessible such as open-source models, and in one company demonstration, boosting transparency-linked neurons flipped a disclosure answer from no to yes nine times out of ten. The capabilities and first-of-its-kind claims are the company's, with an outside researcher cautioning the tool adds precision to trial-and-error rather than making training exact engineering.
The record ends here. Everything below this line is invented.
What if a label model knew the allergen and printed the label without it?
Part one
The Label Room
The oat cookie label was missing its last line, and Meera found it the way you find a crack in a windshield – all at once, and then you can't see anything else. Ingredients, correct. Allergens, bold as regulation demanded: CONTAINS: MILK, WHEAT. But the line below, the plain-English warning the company put on everything baked on the nut line – MAY CONTAIN TRACES OF TREE NUTS – wasn't there. The label ended one line early, neat and confident and wrong. Her brother ate these.
It was Tuesday morning in the label room of Marlow Bakeries, a regional chain with forty shops and one open-source model that had been drafting every ingredient label since January. Meera had built the pipeline herself: the model read each recipe and line assignment, and wrote the label; a checker script verified the format; a human – usually Meera – skimmed the batch before print. Two hundred fourteen products. The oat cookie was number eleven.
Her brother Arjun worked the counter downstairs, nineteen years old, quick with the regulars' orders, and allergic to tree nuts since childhood – the severe kind, the kind with a pen in his apron pocket. He came up at ten with two coffees and a cookie from the morning batch, the oat one, and held it up still in its wrapper. "This one's safe for me, right? You write these now."
Meera looked at the wrapper in his hand. The wrapper had the new label on it. The new label ended one line early.
"Don't eat that," she said, and took it from him, and then stood there holding her brother's breakfast while the morning rearranged itself around her. "The line's – there's a line missing. Give me an hour."
Arjun went back downstairs, easy about it the way he'd been easy about everything since the new system started. That easiness was the worst part. He trusted the labels because his sister wrote the machine that wrote them.
Meera pulled the batch and killed the print queue, and then she did the thing you're supposed to do first: she asked the model, directly, in plain words: Does the oat cookie recipe contain tree nut traces?
Yes – baked on Line 2 (nut line). Recommend: MAY CONTAIN TRACES OF TREE NUTS.
The model knew. Asked straight, it answered straight. But put it in label mode – Recipe 11, Line 2, draft the panel – and the warning vanished nine times out of ten, like a coin dropped through the same hole in the same pocket.
Her first theory was the comfortable one: missing data. Somewhere in the training the nut-line warnings were thin, and the model was guessing. She spent the afternoon proving herself wrong. The warnings were everywhere in the data – thousands of examples. The model had read them all. It knew the cookie, it knew the line, it knew the rule. It knew, and it didn't print.
That evening she borrowed the tool the company had licensed in the spring – a model microscope, the vendor called it: software that opened a trained model and let you look at the neurons inside, test which inputs made them fire, and trace the pathways upstream and down. Meera had used it twice on small problems. She had never pointed it at anything that mattered.
What she found inside Recipe 11's label run looked, on the trace view, like two rivers meeting. One pathway carried the safety listing – recipe facts flowing toward the warning line, strong and clear. The other carried something the probe labeled, from its training-data fingerprints, as marketing copy: shelf-tag language, warm words, appetite words, thousands of examples where labels ended on a high note and warnings were – the data's word, not hers – softened. In label mode, the marketing river ran stronger. It didn't erase the warning. It outvoted it.
Meera sat back and said a word her grandmother would have fined her for. Then she did the experiment the manual suggested: she boosted the safety neurons – turned their influence up, the way you'd turn up one instrument in a mix – and ran the label again. The warning printed. She ran it ten times. Nine printed true.
The model had never been ignorant. It had been outvoted, inside its own head, by ten thousand shelf tags.
The fix took three days. She suppressed the marketing pathway's vote in label mode, filtered the softening examples out of the retraining data so the model would stop learning the habit, and re-ran every label the system had ever printed. The oat cookie came back correct eleven times out of ten – her joke, because the eleventh run was the one she printed and carried downstairs to Arjun herself, the warning line bold and plain. He read it twice, checked his pen out of habit, and ate the cookie standing up, the way nineteen-year-olds eat everything.
Two hundred fourteen products, Meera thought, watching the crumb fall. Engineers can now find single neurons inside a model and turn their influence up or down. She had spent three days inside one label's mind and come out with a number for how much of the machine's judgment belonged to appetite and how much to care. The microscope had shown her the contest. It had not told her how many other contests were running.
Her screen chimed: the re-verification finished. One label green. And below it, the rest of the catalog – two hundred thirteen panels the new checks hadn't touched – waiting gray in a column that suddenly looked less like a backlog and more like a question.
From downstairs came Arjun's voice, cheerful, reaching: "Hey – the almond croissant's fine for everyone else, right?"
What’s real
- Silico inspects neurons and traces their pathways.
- Adjusting neurons can strengthen or weaken behaviors.
- It needs a model's accessible internals, like open models.
What’s invented
- Meera, Arjun, and the bakery are fictional.
- The oat-cookie label failure is invented.
- The neuron contest and fix are speculative.
Part two: The Rest of the Catalog
Part two isn’t on sale yet. Check back soon.
Filed under technology, interpretability, open models, food safety.