
The rapid advancement of large language models (LLMs) has brought about significant challenges, particularly in the form of hallucinations—errors that persist even in the most sophisticated systems. In an effort to tackle this issue, Probably has successfully raised $9 million in seed funding from Andreessen Horowitz to develop a more reliable method for error detection. According to founder Peter Elias, the company's mission is to eliminate hallucinations and factual inaccuracies before they reach the end user. Their ambitious target is to achieve a remarkable 99.99% accuracy, a standard typically found in deterministic systems but elusive in the realm of AI. This goal requires a fundamental reevaluation of many established principles in AI engineering. Probably's inaugural product is a data science tool designed to deliver rapid insights from intricate datasets. What sets this tool apart is its provision of citations and a transparent audit trail for the results it generates—an increasingly essential feature among modern AI applications. To ensure accuracy in its outputs, the system employs a sophisticated validation mechanism that Elias likens to a “data science mech suit.” This mechanism checks the initial answers produced by the LLM against a deterministic validator, ensuring that only results consistent with the dataset are presented. Notably, the LLM has been trained in conjunction with the validator, allowing the entire system to prioritize quick and precise responses. Elias emphasizes that enhancing the harness engineering can enable less powerful models to perform effectively. He explains, “If you can refine the context enough, the model does not have to work very hard to do the right thing. Basically, it’s an exercise in reducing ambiguity.” This innovative approach allows Probably's tool to function efficiently on significantly smaller AI models. Currently, it operates on a model that is “four classes weaker than the frontier models,” meaning it can run on standard local hardware, such as a desktop computer, rather than requiring expansive data center resources. This capability is particularly advantageous as token costs for AI usage continue to escalate, prompting many users to reconsider their AI expenditures. Elias also envisions that this versatile engine can be adapted for various precision-sensitive applications, including accounting and medical services. He remarks, “I think it’s really interesting that the big AI labs have not even attempted to do this. They’re incentivized not to, because they make money the more times you have to correct the model.”
In recent years, the AI sector has been intensely focused on identifying the most advanced models. While this pursuit re...
Business Insider | Jul 25, 2026, 13:10Science Corp is poised to introduce a revolutionary retina chip in Europe, designed to restore partial vision for indivi...
Business Today | Jul 25, 2026, 01:00
In a groundbreaking move to address the critical issue of renewable energy intermittency, a small town in southern Finla...
CNBC | Jul 25, 2026, 05:35
Vietnam is contemplating a distinctive approach to youth social media regulations, diverging from the more common outrig...
TechCrunch | Jul 24, 2026, 21:25
Prentis, a newly established AI research lab, is making waves in the tech industry as it prepares to raise $100 million ...
TechCrunch | Jul 25, 2026, 24:00