An AI Learns to Read a Detector Before Anyone Labels the Data

A neutrino that interacts at tera-electronvolt energy does not leave a tidy trace. It leaves a bright, tightly packed bundle: a shower core, secondary tracks peeling away from it, several kinds of activity piled on top of one another. Much of it spills out of the instrument before it stops. Physicists at ETH Zurich who work on such events put the problem in unusually stark terms in Nature Machine Intelligence on Sept. 30, 2026. At this energy, they argue, the question is not whether machine learning improves on conventional reconstruction software, but whether any practical analysis is possible without it.
Their test case is FASERCAL, a proposed upgrade that would add an energy-measuring calorimeter to FASER, an experiment already running at the LHC and directly observing collider neutrinos since 2023. That upgrade has not been built. Every event in the study is simulated, generated in software and pushed through a model of an instrument whose main volume alone holds hundreds of thousands of readout cells.
That distinction matters, because simulation is also what makes the rest of the problem awkward. It supplies detector data in effectively unlimited quantity. What it does not supply cheaply is the answer key. Training a model to say which kind of neutrino produced an event, or where the interaction began, requires events whose answers have already been worked out and attached, and in this field that is expensive. Rare channels need their own simulation campaigns, systematic variations multiply the cost, and many quantities have to be traced back through the truth record. Labeled examples, not raw events, are the scarce resource.
Saúl Alonso-Monsalve, André Rubbia and colleagues attack that scarcity by training the model on the events first and the questions second. Their encoder, built to take in a detector whose subsystems report in different formats, is first shown events with three-quarters of their occupied regions hidden and asked to reconstruct what is missing. In a second phase it is asked something else entirely: to say what role each individual cell plays. Is this deposit a real one or an artifact of reconstruction? Does it belong to a particle from the original interaction or to something produced downstream? Was it left by an electron or photon, by a muon, or by the heavier particles made of quarks? Both jobs draw on what the simulation already knows about each cell, rather than on the event-level answers the model is eventually asked for.

The payoff shows up when labeled data runs short. The team compared three versions of one architecture: one started from random values and trained from scratch, one pretrained only to fill in the masked regions, and one pretrained on both objectives. Given about a thousand labeled events, the fully pretrained version scored 0.818 on the authors' flavor-identification measure, a separation score on which 1.0 would be perfect. The from-scratch model did not reach that level until it had roughly ten thousand labeled events, where it scored 0.807.
The gap is 0.011, and each score is a mean over three training runs, so the careful verb is the one the authors chose: the pretrained encoder "matches the flavour-classification performance of scratch training with an order of magnitude more data." Masked reconstruction alone did not get there. At the same budget it fell short of what scratch training managed with ten times the labels, and it was the second objective, the one that asks what each cell is, that carried the model over.

The order-of-magnitude saving belongs to one task at one label budget, and the paper does not spread it further. Identifying events that contain a charm quark improves as well, but there the pretrained model matches scratch training on roughly three times more data, not ten. At the smallest budgets of all, pretraining is not reliably better for continuous quantities such as energy and momentum.
Further tests ask how portable the learned representation is. Reused on public datasets from instruments of completely different design, a fine-grained plastic scintillator and a liquid-argon detector, the pretrained encoder beat training those tasks from scratch. On one public benchmark it edged past the best previously published numbers. Then the harder test. The same trained models were re-run, without any retraining, on events built by a different simulation of how neutrinos strike atomic nuclei and matched event by event against the originals. Flavor identification barely moved. Charm identification fell from 0.864 to 0.721, which the authors read as "task-dependent generator robustness, not generator independence."
The abstract closes by saying the authors use "foundation-style" in a restricted sense "rather than claiming a completed general-purpose detector foundation model," and the discussion repeats it. No language model is involved anywhere in this work, and the term is doing narrower duty than it usually does in AI. It means one reusable encoder, trained without the downstream answers, adapted to several tasks and reused outside the detector it learned on.
The authors themselves set out what would take this from a promising simulation study to something an experiment could rely on. Broader generator checks, flux variations, detector-response mismodeling and calibration against real experimental data all "remain necessary before deployment." This is one group's unreplicated result, and it is the same group that proposed the detector and wrote the simulation framework the method is graded on.
Sources
- Nature Machine IntelligencePeer-reviewed
