We value your privacy

    We use cookies to understand how you interact with our website to improve your experience. By accepting, you agree to our use of these cookies. You can always change your mind later.

    Back to Insights
    IFLAI Research

    The Missing Dataset: Failed Experiments

    IFLAI Research
    August 2, 2026
    14 min read

    Every laboratory has two datasets.

    The first one is visible. It is the clean dataset: the successful experiments, the polished figures, the representative images, the assays that behaved, the synthesis routes that worked, the samples that passed quality control, the results that were good enough to publish, report, or send forward.

    The second one is much larger, and usually much less organized.

    It lives in old folders, lab notebooks, failed runs, rejected plates, discarded fields of view, weird spectra, missing metadata, half-finished analyses, and the memory of the person who knows that “this protocol always fails when the cells are too dense” or “that microscope starts drifting after long overnight runs.”

    That second dataset is mostly invisible. And for AI, that is a problem.

    Because if a model only learns from the experiments that worked, it never learns the shape of failure. It learns the polished version of science, not the real one. It learns what a successful result looks like, but not how close that result was to noise, artefact, drift, contamination, batch effects, poor focus, weak staining, bad segmentation, or a protocol that only works under conditions nobody wrote down.

    In science, the missing dataset is not more positive examples. It is failed experiments.

    The missing dataset: scientific AI often learns from successful experiments while the richer record of failed, noisy, or rejected data remains unused.

    Failure is not the opposite of data

    A failed experiment feels like waste when the goal is a clean result. It consumed time, reagents, instrument access, and human attention without producing the answer we wanted. So it is natural that failed experiments are often pushed aside. They are hard to write up, hard to organize, and hard to turn into a compelling story.

    But from the perspective of learning, failure is not empty.

    A failed experiment tells you where the boundary is. It tells you what conditions do not work, which measurements are unreliable, which artefacts are common, which inputs confuse the system, which protocols are fragile, and which assumptions break when they leave the ideal setting. That information is incredibly valuable.

    In chemistry and materials science, this has been clear for years. A 2016 Nature paper showed that “dark” reactions-failed or unsuccessful hydrothermal syntheses recovered from archived lab notebooks-could be used to train machine-learning models for materials discovery. The key point was not that failure is romantic. It was that failed reactions contain information about the conditions under which products do not form, and those boundaries help a model make better decisions [1].

    More recently, work in chemistry has made the same argument from another direction: low-yield or no-yield reactions are rarely included in the published literature, even though they are exactly the kind of negative data machine-learning models need to understand chemical reactivity more realistically [2]. A 2025 Science Advances study also showed that negative chemical reaction data can improve language-model-based reaction outcome prediction, especially when successful examples are limited [3].

    The lesson extends far beyond chemistry. A model that never sees failure becomes overconfident. It learns the success distribution, but not the edge of the map.

    The published world is biased toward success

    This is not only a data-management problem. It is a cultural problem.

    Science has long had a publication bias toward positive results. Negative results are harder to publish, less likely to be written up, and often treated as less valuable, even when they are carefully designed and informative. A recent position paper in machine learning makes the same point for AI research itself: methods that do not beat the state of the art are often abandoned, even though negative results can reveal failure modes, reduce wasted effort, and make the field more honest about what actually works [4].

    The result is a distorted view of reality. The scientific literature tells us what succeeded often enough to be reported. It is much worse at telling us what was tried and failed. For human readers, that is already a problem. For AI systems, it is worse, because models trained on published or curated success cases can inherit the bias of the record itself.

    They may learn that certain outcomes are more reliable than they really are, because the failures were never visible. They may learn that a protocol is robust, because the fragile attempts disappeared. They may learn that a feature is meaningful, because the artefacts that mimic it were filtered out before training.

    This is one reason scientific AI cannot simply rely on bigger models and more polished datasets. If the underlying data record is biased toward success, scaling the model only scales the bias.

    Data-centric AI has already shifted attention from model design alone toward the quality, quantity, maintenance, and lifecycle of data. That shift matters because many deployed AI failures are not caused by a lack of parameters. They are caused by bad labels, missing metadata, hidden leakage, poor sampling, distribution shift, and data that do not reflect the conditions where the model will actually be used [5].

    Failed experiments belong in that discussion. Not as messy leftovers, but as part of the data lifecycle.

    In imaging, failure is everywhere

    In high-throughput imaging, failed data are not unusual. They are part of the workflow.

    A plate may have edge effects. A field of view may be out of focus. Cells may be too sparse, too dense, dead, drifting, poorly stained, overexposed, underexposed, contaminated, segmented incorrectly, or simply uninformative. A treatment may produce no visible phenotype. A batch may look different from the previous one for reasons that have nothing to do with biology. A model may classify an artefact with complete confidence because it has never been trained to know that this artefact exists.

    Most of this data is filtered away before the “real” analysis begins. That makes sense if the only goal is to produce a clean figure. It makes less sense if the goal is to build reliable AI.

    A useful imaging model should know what good data look like, but it should also know what bad data look like. It should know when a field of view is not worth analyzing. It should know when an image is outside the distribution it was trained on. It should know when a segmentation mask is likely wrong, when a phenotype score is unreliable, and when the right answer is not a classification, but a warning.

    This is where failed experiments become a form of supervision.

    A rejected image can teach quality control. A bad batch can teach robustness. A failed annotation can reveal ambiguity. An outlier can become an active-learning target. A negative phenotype can be just as useful as a positive one, because it helps define the boundary of the biological response.

    The mistake is thinking that only successful measurements contain signal. Sometimes the failure is the signal.

    Failed imaging data are not just noise. Out-of-focus images, artefacts, dead regions, batch effects, and negative phenotypes can teach AI systems where reliable analysis begins and ends.

    The most useful models know when not to answer

    One of the most dangerous properties of an AI system is confidence in the wrong setting.

    A model trained only on clean examples may behave well as long as reality keeps looking like the training set. But real laboratories are not static. Instruments drift. Users change. Protocols evolve. Samples vary. New artefacts appear. The model eventually sees something it was not trained to handle.

    The question is whether it notices.

    Out-of-distribution detection is the field concerned with exactly this problem: identifying when a model is facing inputs outside its intended knowledge boundaries. Carnegie Mellon’s Software Engineering Institute describes this as a critical capability for trustworthy AI, because systems should be able to recognize unusual inputs rather than force-fitting them into known categories [6].

    For laboratory AI, that capability is not optional.

    If an imaging model sees a contaminated sample, it should not confidently return a biological interpretation. If an assay drifts outside its validated range, it should not pretend that the output is comparable to last week’s data. If a cell state is unfamiliar, it should not hide uncertainty behind a smooth probability score.

    It should say: this looks different.

    And to learn that, the model needs examples of different. Failed experiments, poor-quality runs, artefacts, rejected images, and strange outliers are not just operational noise. They are the training data for knowing when not to trust the output.

    Closed-loop labs need a memory of failure

    This becomes even more important as laboratories move toward automation and agentic workflows.

    In a passive workflow, failed data are often cleaned away after the fact. In a closed-loop workflow, failure can immediately change what the system does next. If a field of view is uninformative, the system can move on. If an instrument begins to drift, it can pause. If a region is uncertain, it can collect more evidence. If a batch looks strange, it can request human review.

    That only works if failure is visible to the system.

    Recent autonomous laboratory work makes this point in practice. In Berkeley’s A-Lab, an autonomous materials synthesis platform, only 30% of the 353 synthesis recipes tested produced their intended targets, even though the system ultimately synthesized 36 of 57 target materials. Crucially, the failures were not meaningless. The system used observations from failed synthesis attempts to infer unfavorable reaction pathways, reduce the experimental search space, and improve later synthesis routes through active learning [7].

    That is the mindset scientific AI needs. Not every failed run is a disaster. Some failed runs are map-making. They tell the system which parts of the experimental space are blocked, unstable, redundant, or not worth exploring again.

    A similar logic is appearing in laboratory quality control. Recent work on automated HPLC experiments in a cloud laboratory used active learning and human-in-the-loop annotation to detect air bubble contamination, a common experimental anomaly that usually requires expert chemists to identify. The resulting system was designed not just to classify data, but to flag affected runs in real time and support continuous quality control in automated laboratories [8].

    That is a different view of failure. Failure is no longer something to hide at the end of the workflow. It becomes something the system watches for, learns from, and uses to improve the next action.

    The lab notebook should become a learning system

    A lot of scientific AI still treats the dataset as a static object: collect data, clean data, train model, evaluate model, deploy model.

    But real laboratories are continuous. They accumulate experience. People learn which samples are fragile, which instruments are temperamental, which artefacts recur, which protocols are robust, and which conditions are not worth repeating. Much of that knowledge never enters the dataset. It remains informal, personal, and easy to lose.

    The next generation of scientific AI needs a better memory.

    Not just storage. Not just raw images. Not just successful results. A real laboratory memory should capture context: what was tried, what worked, what failed, what was rejected, why it was rejected, what changed, which batch it came from, which instrument produced it, which operator ran it, which model version analyzed it, and what the human expert decided afterwards.

    That kind of record would change how models are trained.

    A model could learn from positive examples, negative examples, artefacts, borderline cases, uncertain cases, and human corrections. It could build a richer understanding of the experimental world, not just the clean subset that survived preprocessing. It could ask better questions because it would remember what had already failed.

    This is especially important for active learning. If the model is allowed to request new labels or new measurements, it should not only ask for examples that look promising. It should ask for examples that reduce uncertainty, clarify a boundary, explain a failure mode, or reveal whether an apparent pattern is real.

    In that sense, the best future datasets may not be the largest ones. They may be the most honest ones.

    A better laboratory memory captures successful results, failed runs, artefacts, uncertainty, metadata, and expert corrections as part of one learning system.

    Not all failure data are useful

    There is an important caveat. Failed experiments are valuable only if they are recorded well enough to be interpreted. A folder of unlabeled bad images is not automatically a useful dataset. A failed assay with no metadata may be impossible to distinguish from a file-management problem. A negative result without enough context can mislead as easily as it can help.

    The goal is not to keep everything blindly. The goal is to make failure legible.

    That means recording why data were rejected. Was the image out of focus? Was the sample dead? Was the label ambiguous? Was the phenotype absent? Was the instrument unstable? Was the protocol wrong? Was the model uncertain? Was the human unsure? Was the result genuinely negative, or was the experiment technically invalid?

    These distinctions matter. A true negative result says something about the system. A technical failure says something about the measurement. An artefact says something about the workflow. An outlier says something about the boundary of the model’s knowledge. Lumping all of them together as “bad data” wastes the very information that makes them useful.

    For AI, the label “failed” is not enough. The reason for failure is the data.

    What this means for IFLAI

    At IFLAI, we think this is one of the overlooked foundations of useful scientific AI.

    Better models matter. Better architectures matter. Better representation learning matters. But none of it is enough if the model only sees the cleaned-up version of the experiment. To make AI systems that are robust in real laboratories, we need to train and evaluate them on the messy boundary between signal and failure.

    That means building workflows where rejected images, failed runs, uncertain predictions, annotation disagreements, batch artefacts, and quality-control decisions are not thrown away by default. It means treating them as structured feedback. It means designing AI systems that can learn not only what success looks like, but where success stops.

    This is also where data-efficient and agentic AI come together.

    A data-efficient model should not ask for more data blindly. It should ask for the right data. An agentic system should not only collect more measurements. It should collect measurements that clarify uncertainty, expose failure modes, and make the next experiment smarter.

    That is the real value of failed experiments. They teach the model what the world refuses to do.

    The missing dataset

    Science does not progress only by confirming what works. It also progresses by learning what does not.

    AI should be the same.

    If we want scientific AI to become reliable outside curated benchmarks and polished demonstrations, we need to stop pretending that failed experiments are just waste. They are part of the map. They define the boundaries, expose the artefacts, reveal the weak points, and teach models when to be careful.

    The laboratories that capture this information well will have a different kind of advantage. Not just more data, but better memory. Not just successful examples, but a deeper understanding of the experimental landscape.

    The missing dataset has been there all along. It is everything that did not work. And it may be exactly what AI needs to become useful in the real world.


    References

    • [1] Nature, "Machine-learning-assisted materials discovery using failed experiments" (2016)
    • [2] Journal of Organic Chemistry, "Negative Data in Data Sets for Machine Learning Training" (2023)
    • [3] Science Advances, "Negative chemical data boosts language models in reaction outcome prediction" (2025)
    • [4] ICML, "Position: Embracing Negative Results in Machine Learning" (2024)
    • [5] arXiv, "Data-centric Artificial Intelligence: A Survey" (2023)
    • [6] Carnegie Mellon Software Engineering Institute, "Out-of-Distribution Detection: Knowing When AI Doesn’t Know" (2025)
    • [7] Nature, "An autonomous laboratory for the accelerated synthesis of inorganic materials" (2023)
    • [8] Digital Discovery, "Machine learning anomaly detection of automated HPLC experiments in the cloud laboratory" (2025)