Zoox has published the reasoning behind its claim to be safer than human drivers. The autonomous vehicle company, owned by Amazon, released a detailed account of its safety case in mid-April, only three weeks after recalling its entire fleet of 105 robotaxis. The document explains how Zoox calculates risk and why it believes its vehicles can operate more safely than a human driver, but it leaves some of the most important numbers blank.
One number to rule them all
Everything in Zoox's safety case reduces to one number: the predicted rate of collision, injury, and fatality events, expressed as miles per event. This is a single estimate that adds together risk from the driving software, from the vehicle itself, and from fleet operations. The result functions as a gate. Before any safety-relevant software release, hardware change, or revision to operating procedure, Zoox updates the safety case and checks that the combined figure still clears its target.
The use of a single aggregate metric is unusual in the autonomous vehicle industry. Some competitors publish separate metrics for disengagements, collisions, and near-misses. Zoox's approach tries to create a single comparable figure that can be weighed against human performance. This makes the safety case easier to evaluate in principle, but also easier to critique if the underlying assumptions are not transparent.
A benchmark built from human data
Zoox anchors its target in human driving data. The company builds its benchmark from the National Highway Traffic Safety Administration's crash sampling and fatality reporting systems, as well as two Federal Highway Administration datasets. It parses those datasets by road speed and weights them to match the mix of roads its robotaxis actually use. This weighting is important because a fleet operating mostly on low-speed urban streets cannot be compared fairly with a national average that includes highways and rural roads.
By matching the benchmark to its own operating domain, Zoox aims to make the comparison more meaningful. If its robotaxis drive in dense city environments, the human crash rate on those same types of roads becomes the relevant baseline. The company says its target is anchored in this human data, but it does not publish the numerical benchmark itself. That omission makes independent verification difficult.
What remains unpublished
Even though the methodology is described in detail, Zoox does not publish the answer. It sets its target by comparison to the human benchmark and says it aims to be significantly safer than a human driver, but it never defines how much safer counts as significant. It also does not give the figure that its fleet currently reaches. This leaves a critical gap: the public can see how the number is calculated, but not what the number is.
This lack of specificity could be intentional. Safety data in the autonomous vehicle industry is often treated as a competitive advantage, and liability concerns may discourage companies from publishing exact figures. But for regulators and the public, the absence of a concrete number makes it harder to assess whether Zoox is actually meeting its own safety goals. The framework is essentially a promise to calculate risk rigorously, without a clear way to verify the outcome.
The engineering behind the estimate
The engineering described in the safety case is specific enough to argue with. Zoox names the hazard analyses it runs, follows the ISO 26262 functional safety standard with integrity ratings on the platform, and uses simulation that deliberately searches for the conditions where a collision is most likely. Those simulation results are later weighted by real fleet exposure to produce a realistic risk estimate.
This combination of directed simulation and real-world weighting is more sophisticated than a simple log of miles driven. It allows Zoox to explore rare edge cases that may not appear often in normal testing. The company also says it uses separate integrity ratings on different parts of the platform, which means critical systems are held to a higher standard than non-critical ones. This aligns with established automotive safety practice.
A second pair of eyes
Some aspects of the Zoox system are genuinely unusual. The company says it runs a separate collision checker with its own perception system. This independent checker can veto a trajectory that the main planning system has proposed. This is a more robust approach than relying on a single perception stack to both see and act. If one perception system fails to detect an obstacle, the other may still catch it.
Zoox also concedes that a likelihood-based metric cannot capture rare avoidance scenarios. A statistical model may show that a particular situation is unlikely to occur, but if it does occur, the consequences could be severe. To address this, Zoox keeps a test set where the robotaxi must at least match a competent human driver. This test set is separate from the statistical risk model and serves as a safety net for edge cases that probability alone might miss.
Humans in the loop
Another notable feature of the Zoox framework is the inclusion of remote staff inside the model rather than outside it. The company employs TeleGuidance tacticians who never drive. They offer route guidance and high-level suggestions while the vehicle itself remains responsible for all driving decisions. The risk of these tacticians making a mistake, or their tools failing, is priced into the same safety estimate.
This is a different philosophy from companies that use remote operators as a fallback with full driving authority. Zoox argues that keeping human responsibilities limited reduces the cognitive burden on remote staff and makes their behavior more predictable. But the framework explicitly accounts for human error, acknowledging that even a limited role carries some risk.
A framework with teeth
The safety case also reserves the right to restrict, pause, or ground the fleet. This is not a theoretical power. Three weeks before publishing the framework, Zoox recalled all 105 of its robotaxis after one failed to detect heavy smoke and drove into an active fire scene in Las Vegas. That incident involved a vehicle continuing into an area where first responders were operating, a serious failure that could have endangered emergency workers.
The recall was the fourth software recall for Zoox in roughly 13 months. It followed regulatory pressure from the National Highway Traffic Safety Administration, which had demanded fixes for vehicles interfering with first responders. NHTSA had logged 123 collisions involving Zoox vehicles in autonomous mode as of March. Although many of those collisions may have been minor, the volume underscores the challenges of operating a robotaxi service in real-world conditions.
The repeated recalls raise questions about the maturity of Zoox's software. The company has been developing autonomous vehicles since 2014, but its commercial deployment is still early. Each recall is a sign that the safety case is working as intended, catching problems before they cause more severe accidents, but the frequency also suggests that the system is still evolving significantly.
The mileage gap
The analytical approach exists because the mileage does not. Zoox has around three million autonomous miles, a substantial figure for a young company but far behind its main competitor. Waymo passed a hundred million autonomous miles more than a year ago. This gap matters for both engineering and public trust. More miles generally mean more exposure to rare events and more data to validate safety claims.
Zoox now has a paid service to protect while it closes that gap. The company launched its robotaxi service in Las Vegas in 2023, after years of testing in San Francisco and other cities. It offers fares to the public, which means its safety decisions directly affect paying passengers. The publication of the safety framework can be seen as an effort to build confidence among riders and regulators, even as the company acknowledges that its most important metrics remain secret.
For now, Zoox is asking the public to accept a detailed methodology without a final score. That may be a reasonable request in a complex engineering field, but it will likely face continued scrutiny from regulators who have already forced multiple recalls. The path to proving safety is long, and Zoox is still writing it, one miles-per-event calculation at a time.
Source: TNW | Amazon News