Showing posts with label road testing. Show all posts
Showing posts with label road testing. Show all posts

Regulating Automated Vehicles with Human Drivers

Summary 

Regulatory oversight of automated vehicle operation on public roads is being gamed by the vehicle automation industry via two approaches: (1) promoting SAE J3016, which is explicitly not a safety standard, as the basis for safety regulation, and (2) using the "Level 2 loophole" to deploy autonomous test platforms while evading regulatory oversight. Regulators are coming to understand they need to do something to reign in the reckless driving and other safety issues that are putting their constituents at risk. We propose a regulatory approach to deal with this situation that involves a clear distinction between production "cruise control" style automation that can be subject to conventional regulatory oversight vs. test platforms that should be regulated via SAE J3018 use for testing operational safety.

Video showing Tesla FSD beta tester unsafely turning into oncoming traffic.

Do not use SAE J3016 in regulations

The SAE J3016 standards document has been promoted by the automotive industry for use in regulations, and in fact is the basis for regulations and policies at the US federal, state, and municipal levels. However, it is fundamentally unsuitable for the job. The issues with using SAE J3016 for regulations are many, so we provide a brief summary. (More detail can be found in Section V.A of our SSRN paper.)

SAE J3016 contains two different types of information. The first is a definition of terminology for automated vehicles, which is not really the problem, and in general could be suitable for regulatory use. The second is a definition of the infamous SAE Levels, which are highly problematic for at least the following reasons:

In practice, one of two big issues is the "Level 2 Loophole" in which a company might claim that the fact there is a safety driver makes its system Level 2, while insisting it does not intend to ever release that same automated driving feature as a higher level feature. This could be readily gamed by, for example, saying that Feature X, which is in fact a prototype fully automated driving system, is Level 2 at first. When the company feels that the prototype is fully mature, it could simply rebrand it Feature Y, slap on a Level 4 designation, and proceed to sell that feature without ever having applied for a Level 4 testing permit. We argue that this is essentially what Tesla is doing with its FSD "beta" program that has, among other things, yielded numerous social media videos of reckless driving despite claims that its elite "beta test" drivers are selected to be safe (e.g., failure to stop at stop signs, failure to stop at red traffic signals, driving in opposing direction traffic lanes).

The second big practical issue is that J3016 is not intended to be a safety standard, but is being used as such in regulations. This is making regulations more complex than they need to be, stretching the limits of the esoteric technical expertise in AVs required of regulatory agencies, especially for municipalities. This is combined with the AV industry promoting a series of myths as part of a campaign to deter regulator effectiveness at protecting constituents from potential safety issues. The net result is that most regulations do not actually address the core safety issues related to on-road testing of this immature technology, in large part because they aren't really sure how to do that.

For road testing safety purposes, regulators should focus on both the operational concept and technology maturity of the vehicle being operated rather than on what might eventually be built as a product. In other words "design intent" isn't relevant to the risk being presented to road users when a test vehicle veers into opposing traffic. Avoiding crashes is the goal, not parsing overly-complex engineering taxonomies.

The solution is to reject SAE J3016 levels as a basis for regulation, instead favoring other industry standards that are actually intended to be relevant to safety. (Again, using J3016 for terminology is OK if the terms are relevant, but not the level definitions.)

Four Regulatory Categories

We propose four regulatory categories, with details to follow:

  • Non-automated vehicles: These are vehicles that DO NOT control steering on a sustained basis in any operational mode. They might have adaptive speed control, automatic emergency braking, and active safety features that temporarily control steering (e.g., an emergency swerve around obstacles capability, or bumping the steering wheel at lane boundaries to alert the driver).
  • Low automation vehicles: These are vehicles with automation that CAN control steering on a sustained basis (and, in practice, also vehicle speed). They are vehicles that ordinary drivers can operate safely and intuitively along the lines of a "cruise control" system that performs lane keeping in addition to speed control. In particular, they have these characteristics:
    • Can be driven with acceptable safety by an ordinary licensed driver with no special training beyond that required for a non-automated version of the same vehicle type.
    • Includes an effective driver monitoring system (DMS) to ensure adequate driver alertness despite inevitable automation complacency
    • Deters reasonably foreseeable misuse and abuse, especially with regard to DMS and its operational design domain (ODD)
    • Safety-relevant behavioral inadequacies consist of omissive behaviors rather than actively dangerous behavior
    • Safety-relevant issues are both intuitively understood and readily mitigated by driver intervention with conventional vehicle controls (steering wheel, brake pedal)
    • Automation is not capable of executing turns at intersections.
    • Field data monitoring indicates that vehicles remain at least as safe as non-automated vehicles that incorporate comparable active safety features over the vehicle life.
  • Highly automated vehicles: These are vehicles in which a human driver has no responsibility for safe driving. If any person inside the vehicle (or a tele-operator) can be blamed for a driving mishap, it is not a highly automated vehicle. Put simply, it's safe for anyone to go to sleep in these vehicles (including no requirement for a continuous remote safety driver) when in automated operation.
  • Automation test platforms: These are vehicles that have automated steering capability and have a person responsible for driving safety, but do meet one or more of the listed requirements for low automation vehicles. In practical terms, such vehicles tend to be test platforms for capabilities that might someday be highly automated vehicles, but require a human test driver -- either in vehicle or remote -- for operational safety.
Non-automated vehicles can be subject to regulatory requirements for conventional vehicles, and correspond to SAE Levels 0 and 1. We discuss each of the remaining three categories in turn.

Low automation vehicles

The idea of the low automation vehicle is that it is a tame enough version of automation that any licensed driver should be able to handle it. Think of it as "cruise control" that works for both steering and speed. It keeps the car moving down the road, but is quite stupid about what is going on around the car. DMS and ODD enforcement along with mitigation of misuse and abuse are required for operational safety. Required driver training should be no more than trivial familiarization with controls that one would expect, for example, during a car rental transaction at an airport rental lot.

Safety relevant issues should be omissive (vehicle fails to do something) rather than errors of commission (vehicle does the wrong thing). For example, a vehicle might gradually drift out of lane while warning the driver it has lost lane lock, but it should not aggressively turn across a centerline into oncoming traffic. With very low capability automation this should be straightforward (although still technically challenging), because the vehicle isn't trying to do more than drive within its lane. As capabilities increase, this becomes more difficult to design, but dealing with that is up to the companies who want to increase capabilities. We draw a hard line at capability to execute turns at intersections, which is clearly an attempt at high automation capabilities, and is well beyond the spirit of a "cruise control" type system.

An important principle is that human drivers of a production low automation vehicle should not serve as Moral Crumple Zones by being asked to perform beyond civilian driver capabilities to compensate for system shortcomings and work-in-progress system defects. If human drivers are being blamed for failure to compensate for behavior that would be considered defective in a non-automated vehicle (such attempting to turning across opposing traffic for no reason), this is a sign that the vehicle is really a test platform in disguise.

Low automation vehicles could be regulated by holding the vehicles accountable to the same regulations as non-automated vehicles as is done today for Level 2 vehicles. However, the regulatory change would be excluding some vehicles currently called "Level 2" from this category if they don't meet all the listed requirements. In other words, any vehicle not meeting all the listed requirements would require special regulatory handling.

Highly automated vehicles

These are highly automated vehicles for which the driver is not responsible for safety, generally corresponding to SAE Levels 4 and 5.  (As a practical matter, some vehicles that are advertised as Level 3 will end up in this category in practice if they do not hold the driver accountable for crashes when automation is engaged.)

Highly automated vehicles should be regulated by requiring conformance to industry safety standards such as ISO 26262, ISO 21448, and ANSI/UL 4600. This is an approach NHTSA has already proposed, so we recommend states and municipalities simply track that topic for the time being. 

There is a separate issue of how to regulate vehicle testing of these vehicles without a safety driver, but that issue is beyond the scope of this essay. 

Automation test platforms

These are vehicles that need skilled test drivers or remote safety monitoring to operate safely on public roads.  Operation of such vehicles should be done in accordance with SAE J3018, which covers safety driver skills and operational safety procedures, and should also be done under the oversight of a suitable Safety Management System (SMS) such as one based on the AVSC SMS guidelines.

Crashes while automation is turned on are generally attributed to a failure of the safety driver to cope with dangerous vehicle behavior, with dangerous behavior being an expectation for any test platform. (The point of a test platform is to see if there are any defects, which means defects must be expected to manifest during testing.)

In other words, with an automation test platform, safety responsibility primarily rests with the safety driver and test support team, not the automation. Test organizations should convince regulators that testing will overall present an acceptably low risk to other road users. Among other things, this will require that safety drivers be specifically trained to handle the risks of testing, which differ significantly from the risks of normal driving. For example, use of retail car customers who have had no special training per the requirements of SAE J3018 and who are conducting testing without the benefit of an appropriate SMS framework should be considered unreasonably risky.

This category covers all vehicles currently said to be Level 4/5 test vehicles, and also any other Level 2 or Level 3 vehicles that make demands on driver attention and reaction capabilities that are excessive for drivers without special tester training.

Regulating automated test platforms should concentrate on driver safety, per my State/Municipal DOT regulatory playbook. This includes specifically requiring compliance with practices in SAE J3018 and having an SMS that is at least as strong as the one discussed in the AVSC SMS guidelines.

Wrap-up

Automated vehicles regulatory data reporting at the municipal and state levels should concentrate on collecting mishap data to ensure that the driver+vehicle combination is acceptably safe. A high rate of crashes indicates that either the drivers aren't trained well enough, or the vehicle is defective. Which way you look at it depends on whether you're a state/municipal government or the US government, and whether the vehicle is a test platform or not. But the reality is that if drivers have trouble driving the vehicles, you need to do something to fix that situation before there is a severe injury or fatality on your watch.

The content in this essay is an informal summary of the content in Section V of: Widen, W. & Koopman, P., "Autonomous Vehicle Regulation and Trust" SSRN, Nov. 22, 2021. In case of doubt or ambiguity, that SSRN publication should be consulted for more comprehensive treatment.

-----
Philip Koopman is an associate professor at Carnegie Mellon University specializing in autonomous vehicle safety. He is on the voting committees for the industry standards mentioned. Regulators are welcome to contact him for support.


Autonomous Vehicle Testing Guidance for State & City DOTs

Once in a while I'm contacted by a city or state Department of Transportation (DOT) to provide advice on safety for "self-driving" car testing. (Generally that means public road testing of SAE Level 3-5 vehicles that are intended for eventual deployment as automated or autonomous capable vehicles,)

The good news is that industry standards are maturing. Rather than having to create their own guidelines and requirements as they have in the past, DOTs now have the option of primarily relying upon having AV testers conform to industry-created guidelines and consensus standards.

And ... in September 2021 NYC DOT blazed a trail by requiring the self-driving car industry to conform to their own industry consensus testing safety standard (J3018). Kudos to NYC DOT!  (check it out here (link); more on that in the details below.

The #1 important thing to keep in mind is that testing safety is not about the automation technology -- it is about the ability of the human safety driver to monitor and intervene when needed to make safety. The technology is going to fail, because the point of testing is to find surprise failures. If a failure of technology causes a fatality, then most likely the testing wasn't being done safely enough. It is essential that human safety drivers be skilled and attentive enough to prevent loss events when such failures inevitably occur.

The short version is that DOTs should:

  1. Follow the AAMVA road testing guidelines plus some additional key practices.
  2. Define how safe testing should be when considering the safety driver + vehicle system as a whole.
  3. Ask testers for conformance to SAE J3018 for road testing.
  4. Ask testers to have a credible Safety Management System (SMS) approach, including a testing plan.
  5. Ask testers to provide metrics that show that their testing is safe (not just a promise up front, but also periodic testing safety metrics as they operate).  Don't get distracted by measuring the maturity of the technology they are testing -- it's all about the safety driver ability to intervene when something goes wrong.
  6. If testing takes place with a safety driver in a chase vehicle or remote, ask for conformance to safety standards for the mechanisms required to ensure safety (e.g., per ISO 26262), but otherwise conforming to SAE J3018 for the training and protocols.
  7. If testing takes place without continuously monitoring safety driver, ask for conformance to industry-consensus safety standards for the autonomous vehicle itself.  If there is no person continuously monitoring and capable of assuring safety, then the safety aspects of the technology have to be done. You shouldn't let vehicles without fully mature safety technology operate without a human safety driver.


Cars driving from cell phone to road


The long version (below) gets pretty detailed, but this is a complicated and nuanced issue. So, here we go...

The AAMVA Guidelines as a starting point:

DOTs should follow applicable AAMVA guidelines with a few additional points.

The American Association of Motor Vehicle Administrators (AAMVA) released the 2nd edition of Safe Testing and Deployment of Vehicles Equipped with Automated Driving Systems Guidelines in September 2020. There is plenty of good information here. However, there are a few areas that require going beyond these guidelines to ensure what might be considered acceptable safety.

My additional recommendations within the scope of these guidelines include:
  • Vehicle manufacturer or testing organization should be required to publish a Voluntary Safety Self Assessment (VSSA) report (see AAMVA Guidelines 3.1.5). That VSSA should address all relevant topics in the NHTSA Automated Driving Safety documents (2.0, 3.0, 4.0). A VSSA does not provide all information required for technical evaluation of safety, but complete lack of a VSSA suggests an unwillingness to provide public transparency.
  • Require statement of areas of intended operation in a manner that does not compromise any claimed secrets as to detailed specifics of tests being conducted. 
    • For example, require reporting of zip codes of where testing is to be conducted.
    • Report speeds at which testing will be conducted (e.g., 25 mph speed limit street testing is much different than Interstate System highway testing).
    • Report other relevant Operational Design Domain factors that will limit testing (e.g., daytime only, in rain, in snow) so that any particularly hazardous environmental testing situations can be considered with regard to public safety.
    • Discuss any unique characteristics of the test area to ensure the tester understands what unique challenges might be presented that someone not from the location might find unusual (e.g., The Pittsburgh Left, parking chairs, cable cars, cattle grids, gator crossings).
  • Require a tester statement that a defined level of technology quality will be confirmed before it is used for public road testing (along with the definition of what that might be). This should include at least:
    • A comprehensive simulation and closed course testing plan should be completed before testing on public roads.
    • All software updates should be subjected to confirmatory closed course testing to ensure no new defects have been introduced before being used in road testing.
    • Any vehicle feature that does not pass closed course testing should not be active during road testing. (In other words, if a feature fails closed course track testing, it shouldn't be operated on public roads.) Public roads should be used to confirm that the vehicle works as expected, not for debugging of known-faulty features.
  • Explanation for why the tester thinks that safety driver training and performance will be sufficient to ensure that test vehicles do not present increased risk to other road users.
The AAMVA guideline scope, while quite useful, is primarily administrative in nature rather than technical. To go beyond this we need to look at engineering standards. (Some of the above points also appear in the following standards and guidance.)

Define how safe is safe enough:

DOTs should define the desired safety outcome, but not how to measure it.

This is perhaps the trickiest point. It's important for the DOT to set the bar for how safe is safe enough. Testers likely have overwhelming financial incentive to get their testing done. Even with the best intentions, the threat of losing funding for lack of progress can loom larger than a possibility of a problem with a testing crash that might (or might not) happen in the future. It seems insufficient in such an environment to simply assume that for-profit organizations will set a safety target that reflects local societal norms.

However, it would be irresponsible for a testing organization to do public road testing without regard for public safety. This means that any (responsible) testing organization will have: a safety goal and analysis before testing starts to predict whether they are likely to reach that safety goal, and metrics collected during testing to ensure that they are meeting their safety goal.

DOTs might not have the technical sophistication to tell testers how to predict safety during testing, nor to know exactly which metrics and associated metric thresholds would be appropriate for a particular test plan. However, the DOT should take responsibility (absent legislation) for making it clear what the testing safety goal should be.

An example might be: road testing operations shall be at least as safe as unimpaired human drivers, taking into account local driving safety statistics and testing environmental conditions. For example, if testing in Pittsburgh only in daytime and dry weather, testers should have a goal of being at least as safe as other Pittsburgh drivers operating in daytime and dry weather (subtracting out drunk and impaired human driver collisions).  That "safer than human" should consider at least fatalities and major injury crashes. Records must be kept of all safety-related metrics, incidents, and loss events.

Some important considerations is that the policy in the preceding paragraph does not tell testers how to predict such safety nor how to measure it on a technical basis. Rather, it is up to the testers to figure this out in their own individual situation. As mentioned earlier, if they don't know how to measure their own safety, they shouldn't be out on public roads doing the testing in the first place.

Could this approach be gamed? Of course it can (as can any approach). However, if the tester goes on record committing to a particular level of safety, it will become evident whether that level of safety has been reached sooner or later based on police reports, if nothing else. Once that happens, historical metrics will show whether the tester was operating in good faith or not.

SAE J3018 for operational safety:

DOTs should ask testers to conform to the industry standard for road testing safety: SAE J3018. 

When AV road testing first started, it was common for testers to claim that they were safe because they had a "safety driver." However, as was tragically demonstrated in the Tempe AZ testing fatality in 2018, not all approaches to safety driving are created equal.  Much more is required. Fortunately, there is an SAE standard that addresses this topic.

SAE J3018_202012 "Safety-Relevant Guidance for On-Road Testing of Prototype Automated Driving System (ADS)-Operated Vehicles" (https://www.sae.org/standards/content/j3018_202012/ -- be sure to get the 2020 revision) provides safety relevant guidance for road testing. It concentrates on guidance for the "in-vehicle fallback test driver" (also known informally as the safety driver).

AV testers should be conform to J3018 to ensure that they are following identified best practices for safety driver training and effectiveness. Endorsing this standard will avoid a DOT having to create their own driver qualification and training requirements.

Taking a deeper look at J3018, it seems a bit light on measuring whether the safety driver is actually providing effective risk mitigation. Rather, it seems to implicitly assume that training will necessarily result in acceptable road testing safety. While training and qualification of safety drivers is essential, it is prudent to also monitor safety driver effectiveness, and testers should be asked to address this issue. Nonetheless, J3018 is an excellent starting point for testing safety. Testers should be doing at least what is in J3018, and probably more.

J3018 does cost money to read, and the free preview is not particularly informative. However, there is a free copy of a precursor document available here: https://avsc.sae-itc.org/principles-01-5471WV-42925L3.html  that will give a flavor of what is involved. That having been said, any DOT guidance or requirement should follow J3018, and not the AVSC precursor document.

In addition to following J3018, the safety-critical mechanisms for testing should be designed to conform to the widely used ISO 26262 functional safety standard. (This is not to say that the entire test vehicle -- which is still a work in progress -- needs to conform to 26262 during testing. Rather, that the "Big Red Button" and any driver takeover functions need to conform to 26262 to make sure that the safety driver can really take over when necessary.)

For cargo vehicles that will deploy without drivers, J3018 can still be used by installing a temporary safety driver seat in the vehicle. Or the autonomy equipment can be mounted on a conventional vehicle in a geometry that mimics the cargo vehicle geometry. When the time comes to deploy without a driver physically in the system, you are really testing an autonomous vehicle with a chase car or remote safety supervisor, covered in a following section on testing without a driver.

Safety Management System (SMS)

DOTs should ask testers to have a Safety Management System in place before testing.

A Safety Management System is a systematic way to manage safety risk for an organization. The roots of SMS approaches come from the aviation industry. The short version is that an SMS helps make sure that you are operationally safe. An important aspect of an SMS is that traditionally it is more about how the people in the company perform tasks and the safety culture rather than the technology itself.

Perhaps the most important overarching finding of the NTSB investigation of the Tempe AV testing fatality was that the lack of an SMS increased the risk of such a bad outcome. To paraphrase the NTSB hearing opening remarks: "you don't have to wait to have a fatal crash before you decide to implement an SMS."  (If you have made it this far in reading this essay, you absolutely must listen to the first 6 minutes of this NTSB hearing https://youtu.be/mSC4Fr3wf0k if you have not already done so.)

The AVSC, a closed-membership industry group, has recently released guidelines for AV testing SMS: https://avsc.sae-itc.org/principle-7-5896VG-46559OG.html  
While these are not at the same level of consensus and review of an SAE issued standard (for example, public comments are not solicited), they do provide industry guidance that is applicable to road testing safety. (I personally have not reviewed these to the degree I have J3018, but expect to do so over time if it is submitted to the SAE ORAD standards committee as J3016 was. So this is not a specific endorsement, but rather an identification of industry-created content that looks likely to be useful.)

DOTs should ask that any AV testing organization to describe their SMS and accompanying safety plan. The tester should explain how such an SMS is comparable to or better than the AVSC guidelines.

Metrics:

DOTs should ask for metrics related to public safety during testing rather than autonomy performance.

It is common to want a standard set of metrics for both test and deployed vehicles. That area is still maturing. While metrics such as number of crashes of various severity classes are fairly straightforward, other predictive metrics such as "disengagements" are problematic for a number of reasons. In particular, each vehicle and each test program has different objectives and different safety architectures. So it will be a while before one-size-fits-all metrics are standardized.

We recommend that any metrics defined be tied to safety procedures and policies rather than the maturity of the technology. Most importantly, it is desirable to find metrics that AV testers cannot claim reveal proprietary information. That means that metrics that measure "how good is the AV" or "how soon to deployment" are likely to be problematic -- and not necessarily that relevant to the crucial question of whether the testing itself that's going to happen right now (and not the AV that might be deployed sometime in the future) presents elevated risk to the public.

I'd argue that the public has a legitimate right to understand whether road users are in the test area are put at increased risk due to AV testing. One way to approach this is to ask the AV tester to respond the following questions:
  • What basis do you have for claiming that your testing will not present increased risk to other road users, including vulnerable road users?
  • What metrics to you plan to collect to ensure that your system is in fact not presenting any such increased risk?
  • What periodic (e.g., monthly) quantitative report can you give us to show that indeed your testing has not increased the risk to other road users?
In general, the strategy should be to ask the AV tester: "Why do you think you're safe" and "How do you plan to measure safety," followed by "How will you know if you're not as safe as you promised you would be?"

If the AV testing can't promise that they will not increase risk to other road users (especially vulnerable road users), then should they be testing on your roads?  If they don't plan to measure and track their actual on-road risk, then do you find their safe testing promise credible? And if they claim that their testing road risk data is proprietary, does that even make sense?

Some example metrics for testing safety (although applicability depends on the specifics of the situation):
  • How often does the built-in driver monitor signal a driver attention issue?  (It won't be zero, but there should be a defined acceptable threshold set by the AV tester which, if exceeded, should cause a process intervention of some sort.)
  • How often does the safety driver make an erroneous intervention, even though there is no crash or other loss event? (In other words, how many near hits are occurring?)
  • How does the AV tester track skill degradation to determine when it is time for a shift change or even refresher training for a safety driver?
Keep in mind that for most companies testing safety is accomplished via test driver supervision, road safety has much more to do with the reliability of the safety drivers than the automation technology itself. So the above metrics have nothing to do with the automation technology, and everything to do with test driver safety -- which is the part that matters for most AV testing safety.

In the end, the metrics should show that the required level of safety is being achieved. They should also be predictive enough that they are likely to indicate any potential problems BEFORE there is a crash.

Testing Without A Driver:

DOTs should ask about safety during communication loss for remote safety driver testing.
DOTs should ask testers to conform to industry automotive safety standards if there is no supervising test driver.

Eventually, organizations will want to test on public roads without a driver. Indeed California has already issued permits for this. In terms of safety, a primary question to ask is how safety is being assured. 

If there is a remote operator involved, then it is important to ensure that any real time data connectivity and sensor information is sufficient to ensure safety. This is a controversial area, and any company promising that, for example, a remote operator can instantly take over operation in the event of a malfunction should be prepared to offer hard data metrics on control latency (delay introduced by the remote communication system), effectiveness of the vehicle detecting its own malfunctions (very difficult to ensure if the system doesn't know it doesn't see a pedestrian for example), and communication link reliability. It is challenging (some would say implausible) to control high speed vehicle operation remotely due to the latencies involved, so a line of sight radio link with a chase car might be required. J3018 practices for the remote operator would still apply. Additionally the equipment used to perform the remote operation should conform to ISO 26262 or other comparable safety standard, which is not typically true for telecommunication equipment. (If loss of signal triggers a vehicle shutdown, then that loss of signal equipment and shutdown mechanism should conform to ISO 26262.)

If there is no remote operator involved, then either the tester should be following issued safety standards or have a safety case to explain why what they are doing is at least as rigorous as what is in those standards. Currently issued and applicable safety standards include: ISO 26262 (functional safety), ISO 21448 (safety of the intended function), and ANSI/UL 4600 (system level safety for autonomous vehicles).

It is worth noting that misinformation has been provided to at least one state DOT regarding ANSI/UL 4600 by industry advocacy groups. (Short version: there is no requirement whatsoever for external assessment in 4600, despite multiple statements to the contrary in a letter sent to a state DOT. Other negative statements tend to be similarly misleading or just plain incorrect.) Any DOT who wants the full story in response to any information they receive criticizing ANSI/UL 4600 is welcome to contact the author of this essay.

Some testers may say they have reasons for not following industry consensus safety standards. If that is the case, ask them what quantitative data they have to demonstrate they are safer than a human driver. If they can't prove to themselves that they are at least as safe as a human driver, why are they operating on public roads? If they say they have the data but it is proprietary, ask what road testing safety data has to do with the secret sauce behind their autonomy.  (Short answer -- it has nothing to do with the secret sauce, but might have to do with concerns that they can't promise safe testing.)

Transparency:

It is common for testers to claim that any attempt to require data reporting, metrics, or other transparency will somehow give away incredibly valuable trade secrets and inhibit innovation. This is utter nonsense. Yet, it seems to be the industry playbook.  For example, during a NYC DOT hearing "about a half-dozen autonomous car makers and their advocates said the proposed rules would turn New York City from an engine of innovation into a backwater that would set back the evolution of the potentially life-saving technology of computer-controlled cars and trucks that can move around without inferior human beings messing everything up."  (https://nyc.streetsblog.org/2021/09/01/self-driving-car-industry-promising-safety-pushes-back-on-dot-plan-to-regulate-testing/)

Often this conversation boils down to testers saying "trust us, we're smart." They may be smart, but decades of experience with safety in other domains has shown that there is no safety without transparency. If they are smart enough to be able to build a car that can drive itself safely on your roads -- without even needing to follow industry standards -- they should also be smart enough to figure out a way to show you data to prove they are safe without revealing major secrets.

For situations in which a safety driver is in the vehicle, let's look at what is required for transparency, which NYC DOT did a good job with (here: https://rules.cityofnewyork.us/wp-content/uploads/2021/08/DOT-Notice-of-Adoption-AV-Rule-FINAL-with-Finding.pdf). The elements they require are:
  • Self-certification that the testing will be safer than a human driver. This is just asking the tester to claim (without producing any proof) that they will test safely. If they're not willing to sign up to that, probably they should not be on public roads.
  • Conform to SAE J3018 and AVSC 00001201911. In other words, this is asking them to follow industry standards and practices for their test driver qualification and testing protocols. This involves ONLY the human test driver and does not place constraints on the automation technology being tested. If they're not willing to sign up to have trained safety drivers and safe testing protocols, probably they should not be on public roads.
  • Submission of a safety plan. This has nothing to do with the automation technology -- it is all about making sure the safety driver can keep the vehicle safe. If they can't explain to the DOT what their plan is to be safe in testing, probably they should not be on public roads.
The key is: you don't need to disclose any autonomous vehicle secret sauce to explain why testing will be safe, because the safety hinges on the human safety driver, not the automation technology.

Other Resources.

Here are some resources that might be useful. While SAE J3016 is widely used for terminology, it is essential to note that it is not (and is not intended to be) a safety standard. Conformance to J3016 Levels has to do with whether you're using an appropriate name for your vehicles, and not whether those vehicles are safe. 

Prof. Philip Koopman is an internationally recognized expert on Autonomous Vehicle (AV) safety whose work in that area spans 25 years. He is also actively involved with AV policy and standards as well as more general embedded system design and software quality. His pioneering research work includes software robustness testing and run time monitoring of autonomous systems to identify how they break and how to fix them. He has extensive experience in software safety and software quality across numerous transportation, industrial, and defense application domains including conventional automotive software and hardware systems. He was the principal technical contributor to the UL 4600 standard for autonomous system safety issued in 2020. He is a faculty member of the Carnegie Mellon University ECE department where he teaches software skills for mission-critical systems. In 2018 he was awarded the highly selective IEEE-SSIT Carl Barus Award for outstanding service in the public interest for his work in promoting automotive computer-based system safety.

Any city or state DOT representative addressing this topic is welcome to contact him via:  koopman@cmu.edu

Updated Sept. 12, 2021.

Road Testing Safety Metrics (Metrics Episode 3)

To build trust, self-driving car companies should be transparent about operational safety metrics for road testing, such as how often the safety driver fails to react to an issue. This won't be perfect, but it ought to be at least as safe as normal human drivers on public roads.

Right now when you see what looks like a self driving car on the road, it’s not really a production self driving car -- it’s a self driving car technology test platform. That usually means that some human someplace in the car remotely, whatever, is keeping an eye on it and making sure things are safe. In turn if you asked the question, “How safe is that car,” you’re not actually asking about the safety of the self driving technology pretty much at all. What you’re asking is whether or not that human is able to properly supervise the safety and do the right thing when something goes wrong. So if you care about the safety of this public on road testing of this maturing technology what you care about is the human safety driver’s performance. To understand that let’s talk about the types of things that safety driver has to do.

The safety driver has to build and maintain situational awareness and know what’s supposed to happen next. The safety driver has to notice something is going wrong and intervene at exactly the right time. While this is happening the safety driver is under some pressure to balance getting operational data with maintaining safety. After all, sitting in the garage with the car turned off doesn’t get the data that they’re out in the public roads to get. As they’re operating the safety driver has to figure out how to react when other drivers do things that are weird, illegal, dangerous, or just plain crazy.  And sometimes it’s the car itself that misbehaves due to a software defect in the self driving technology. When something bad happens, the safety driver has to execute a takeover maneuver and make sure that they don’t make things worse.

Now, that sounds like a lot to do but the fact is that when things are going well it’s an exceptionally boring job. The car just keeps doing the thing it’s supposed to be doing and the driver sits there watching and waiting. Every once in a while something will go wrong and the safety driver has to make sure that not only do they have situational awareness, and that they have an idea of what to do next, but also that they’re not startled into doing the wrong thing. Rolling all that up, supervising an autonomous test platform is exceptionally difficult work and requires the highest of driving skills. And yes, we actually do get to the part where the car has to behave properly, but it has nothing to do with the self driving technology -- it has to do with the emergency override. If the human safety driver wants to override the vehicle then they have to be able to turn off the self driving feature and get control of the vehicle.

It turns out getting that mostly right is straight forward but getting it 100% right takes a lot of care and attention. Taking all this into account, you start realizing that disengagements, which is the measure we use now, is actually the wrong metric. In fact it’s exactly the wrong metric. That's because disengagements aren’t what make you dangerous, the thing that makes you the most dangerous when you’re testing technology is non disengagements: a failure to disengage when you should have. In other words, the hazards that you did not catch with the human safety driver that’s what matters.

Now, measuring something that doesn’t happen is not straightforward but in fact if you want metrics for safety that’s what you have to go after. The starting point isn’t really quite metrics at all but rather processes and what you hear is the more thoughtful companies are starting to emphasize things like good driver qualifications, they have training, they have good reaction times. Driver testing, go out on a test track, make sure the driver can handle surprises and make sure the driver can maintain situational awareness. Yes, the disengagement mechanism has to work. The proverbial big red button needs to work all the time not just most of the time. Fortunately there’s a safety standard ISO 26262 which tells you exactly how to make that kind of system so these companies should be following that standard with their disengagement mechanism.

Now there’s an issue about whether or not training once is enough.  It’s not. You need refreshers at least, and so refresher training is also a good idea. Many of the companies stop with that. But the problem with doing only training is that you think you’re good enough but you don’t know you’re actually good enough unless you take operational data. 

Sure, the driver can pay attention on test day and maintain engagement. But can a safety driver maintain that same peak performance for six months of day in, day out testing that’s really boring because mostly things work pretty well? Humans aren’t perfect, and that means even really good safety drivers aren’t perfect.  That’s just the way it is. So what you really want is not just training but also operational metrics to make sure that the safety drivers are able to be as good as they need to be to reach your safety goal.

Here are some metrics that could work internally for an engineering effort:

  • You could measure the attention of the driver throughout a shift. Maybe you’ll find out, as might be expected, that after many hours the driver has trouble maintaining attention. That wouldn’t be a surprise. Some companies are doing two hour shifts. While for active driving that might be a good shift size, there’s lots of data for several decades suggesting that for supervising autonomy you need shorter shifts, maybe only 30 minutes. So do you have data showing that your two hour shifts are okay? Or maybe you need a shorter shift?  Without data you won’t know that.
  • You might also want to measure the intervention accuracy. Did the driver actually know if their blind spot was clear when they made a lane change, or did they just get lucky when they did a takeover and a lane change? Knowing the fraction of time that drivers do the right thing is important. It’s not going to be 100%, but you just need to make sure it’s good enough to achieve your safety goal. 
  • And every once in a while if the driver tries to do a takeover and the vehicle doesn’t respond you really need to know that that happened.

Now those metrics are somewhat detailed. It might not be the kind of thing that a government agency or the public would really be in a position to process well. So you want some roll up metrics as well. But these are mostly things that have to do with non disengagements rather than disengagements. 

To really deal with road testing safety, the number one concern has got to be times when the vehicle did something dangerous and the safety driver did not do the right compensating action. Your drivers aren’t going to be perfect. Are you planning on just blaming them when there’s a mishap or are you actually taking measurements to find out if you’re hitting performance targets?  So if you have a target for how much the driver has to pay attention and how much is okay to lose attention for a second or two, then you need to know if you’re hitting that target and it’s not going to be perfection.

You probably also want some metrics for rule violations. 

  • For example, running a stop sign even though nothing bad happened, is still a bad thing. The number of times that happens is not going to be zero, but it should be so low that it’s dramatically better than normal human drivers for example.
  • You probably want to measure how often you unreasonably encroach on a bike lane even if there’s no bicyclists there.
  • You might want to measure how often you fail to yield to a pedestrian. Most pedestrians aren’t going to jump out in front of a car to get hit, but the fact that they had to back off and let you pass means you weren’t behaving the way you were supposed to. 
  • Probably the most important thing to measure is near misses (or as some people call them near hits) where the margin was just too small even if you got lucky and nothing bad actually happened. For example, maybe you’re supposed to leave three feet of clearance to a bicyclist. (And at high speeds I would hope you do leave at least that, but if you’re down at two or three miles an hour crawling through a dense urban area and you're two feet away from a bicyclist, maybe that’s okay.) You need a definition of near misses that responds to the particular situation but ultimately you want to know how often you had a near miss, meaning you were in a situation more dangerous than it was supposed to be. If your safety driver catches it, then great. But if your safety driver does not, that means you’re operating taking chances that you didn’t intend to take, and knowing that is super important.

No safety program and no safety metric is going to be perfect. It’s unreasonable to expect self driving car testing to be absolutely perfect with not even a fender bender. Surely mishaps will happen especially when interacting with human drivers. But there has to be some sort of strategy. 

A viable strategy might be that you use above average safety drivers who have to deal with the extra complexity of the safety platforms, but they come out overall to be at least as good as normal human drivers. You’re making a big ask to have these safety drivers be better than normal human drivers and compensate for all the difficulties with this technology. Probably it can be done, but you really need to know you got there instead of training the drivers, sending them out on roads and simply hoping that you’re safe enough. You really want some metrics. 

To build trust companies should be exposing at least the high level rollup metrics to the public. They have nothing to do with the secret sauce behind the self driving car but they have everything to do with the safety of the general public as this technology is being tested on public roads.

For the podcast version of this posting, see: https://archive.org/details/metrics-04-road-testing-metrics

Thanks to podcast producer Jackie Erickson.



Disengagements as a progress metric is a bad idea (Metrics Episode 2)

 We should be worried about road testing safety metrics, not disengagements.

A disengagement happens when the autonomy in a self driving car detects an internal problem, or a human test driver takes over control of a self driving car test platform because of safety concerns. Self driving car developers have to report these disengagements, for example, to California. The apparent rationale for requiring these reports is that all things being equal, disengagements per mile might decrease over time as technology matures. Along those lines, eventually when disengagements reached zero, you might think it’s time to deploy the vehicle without a human test driver. The problem is that this model is much too simplistic and more importantly, not all things are equal.

Let’s start with some basics. Not all miles are equal. If you wanted to game disengagements, you could do so by driving around an empty block in beautiful weather at 4:00 AM with no traffic, no pedestrians, nothing on the road, around and around in circles. You get a lot of miles. You wouldn’t learn much, but you get a lot of miles. That’s not at all the same as, for example, trying to drive across all 446 bridges in Pittsburgh during a blizzard. Those miles are just not the same. Another potential problem is that not all safety drivers are equal. Some safety drivers will be more prone to be cautious and others less cautious. Hopefully there is rigorous driver screening so that the test drivers safety drivers are the right amount of cautious, but in fact, this is still an area the industry’s working on. So even with the best intentions, all disengagements might not be equal.

Now what happens next? After this disengagement data is collected, the metrics get published and that leads the media to trend those published metrics into the great disengagement metric horse race. Pundits opine about which company is in the lead and companies who are ahead say, "yeah, look at our low disengagement rate" and so on. Now it’s hard to blame people for doing this because the developers operate in such great secrecy. That’s really the only progress metric out there, but it’s not a good metric. In fact, it’s probably a harmful metric. A big concern is that using disengagements as a metric provide strong incentives for behavior that make things worse instead of better, especially if you’re being judged in it for progress and maybe your next funding round depends on your disengagement metric.

Here’s a problem: the disengagement metric penalizes companies who tell their safety drivers to be extra safe by being extra quick to disengage. So, that means there’s an incentive to tell drivers to give the vehicles a little more slack, which might or might not be as safe as it should be. Now, I’m not saying that people necessarily do this intentionally, but in a very competitive environment, there’s going to be natural pressure to say, well, if it’s on the borderline, let it go to make our numbers look better and probably it’s safe enough. And people might convince themselves of that, even though they’re operating unsafely. 

Another problem is the metric penalizes companies who are working on difficult operational design domains and incentivizes them to chase easy miles. Now again, I’m not saying companies are doing this on purpose, but certainly the incentive is there. In fact, there are good reasons why a company making excellent progress would actually see their disengagements increase rather than decrease. Maybe the company’s expanded its operational design domain to handle more challenging situations. The week that they decide to start operating in rain, I’d imagine the disengagement rate would go up instead of down.

Another reason is maybe safety driver training has been improved and the policies have been changed to improve road test safety at the expense of increased disengagements. I’d love to see that kind of outcome, but it makes the metric look bad. 

Some companies filter the disengagement to say, okay, well we’re only going to report the disengagements that count and the problem is that’s a two edge sword. Sure, it makes sense not to report planned disengagements. If it’s the end of the testing day and you’re going to take the car back to the garage and you turn off autonomy, surely that disengagement should not count.

But because companies are being judged on disengagements, there’s also incentive to gain them a bit. For example, the driver might take over control and maybe it was a dangerous situation, maybe it wasn’t, but because of the pressure of metrics, the company decides to round down and attributed to something else when in fact probably the car should have been doing better and the disengagement should have at least partially counted. You might end up with under-reporting disengagements and that should be a cause for concern.

Let me give a couple of hypothetical examples of the kind of situations which could lead to this kind of bad outcome. For example, let’s say a company only reports disengagement if an after the fact simulation says yes, they would hit something. And then the car goes by a pedestrian and the safety driver disengages because it looks like it’s going to be kind of close to the pedestrian, they don’t want to take a chance. So far, so good. Now let’s say you would’ve missed the pedestrian by 10 to 20 feet. Okay, fine. That disengagement probably should not count, but what if you’re only going to miss the pedestrian by one inch? Well, you didn’t hit the pedestrian. You could say, well, that one doesn’t count because we didn’t hit anything. But I’m going to say missing a pedestrian by one inch, that one ought to count.

And so without details about how exactly this reporting has been done, we don’t really know what the numbers mean. 

Here’s another hypothetical example. Let’s say a test vehicle runs a red light, but it’s late at night and no one’s around. The driver looks around, the driver’s been told, unfortunately, to make sure the disengagement numbers look good. There’s no cross traffic. The driver says, you know what? I’m just going to let it run the red light because no harm will be done. There’s a situation where the disengagement doesn’t happen, the number looks good, but in that hypothetical scenario, the driver’s been incentivized to do unsafe things. That made up example brings to a head the real issue here. 

Disengagements might be useful input to some parts of an engineering process, but in a hyper competitive market, they provide all the wrong incentives for road test safety.  Really are we worried about progress? 

Do we really want the Departments of Transportation measuring progress of companies? Their job really has to be keeping people safe on the road. And so if the publicly reported data actually provides incentives to do road testing unsafely to make progress look good, that’s a problem. Historically, those kinds of incentives, they lead to a story that ends badly for everyone. 

Well, it’s interesting to know what progress might be made by the industry. But if you really care about safety, the thing you ought to be worried about right now is road testing safety. Now, some of the reporting data actually does help with that. For example, crashes being reported. Sure, that actually directly measures road testing safety. But all the disengagement metrics and all the buzz really doesn’t help road testing safety -- and in fact might be hurting it. It might be putting pressure on the companies to undermine road testing safety just to make their numbers look better.

I would recommend that California and any other government that’s doing this should stop forcing disengagement reporting and instead encourage the reporting of more productive metrics that are about road testing safety, not about the horse race to get a self driving car deployed. This is not a simple ask. This is actually a hard thing to do. But the industry should step up and propose metrics that have to do with road testing safety. That’s going to help them build trust with the public and it’s going to help the government agencies fulfill their responsibility to ensure that the road testing and eventual deployment is done in a safe, responsible manner.

To learn more, we recommend a paper our team published for SAE World Congress titled “Safety Argument Considerations For Road Testing of Autonomous Vehicles.” This paper gives guidance on a safety case for human supervision of road testing.

For the podcast version of this posting, see: https://archive.org/details/metrics-03-disengagement-metrics

Thanks to podcast producer Jackie Erickson.

Number of miles as a self-driving car progress and safety metric (Metrics Episode 1)

Even if you have the best possible safety drivers, every test mile adds some risk. Make sure every mile of road testing is actually doing something important.

You hear people saying that testing for lots and lots of miles must mean some self-driving car company is better than the rest, or at least in the lead in the so-called race to autonomy. But not so fast, there’s more to it than that.

Some companies have millions of miles of road testing experience, and to be sure that’s an impressive accomplishment. If they have that many miles, certainly they have incentive to boast about it in saying, “Look how many miles we have.” And the press often says, “All right, these guys have lots and lots of miles, somehow that must be in they’re ahead.” But miles really doesn’t tell you who’s safer or even necessarily ahead. Miles is a mostly reflection about the resources they have available to deploy a test fleet.

If you have lots of money, you can buy a lot of cars, hire lots of people, and put them out there to rack up the miles. Sure, having lots of resources makes it easier to make progress, but it doesn’t tell you who’s better. In fact, some companies are taking pride in reducing their testing road miles and instead putting those resources into simulation and other sorts of engineering activities other than road testing. It’s hard to believe that if you don’t have any miles or you only have a few miles that you’re really ready to deploy, I get that. And it’s hard to believe that a company without lots of data collection miles is actually seeing all the things that are going to need to deal with when they pick an operational design domain. But nobody will ever get to the billions of representative miles necessary to say anything compelling about expected safety until after they actually deploy their fleet.

Even if you have a lot of miles, not all miles are created equal. Okay, so we have the billion miles -- are they in simulation around the same block? Are they all in sunny weather -- or do they include rain and hail and ice and all those other types of weather conditions you care about? Are they in a place with wide roads and no pedestrians or are they in a chaotic urban center? Oh, was that chaotic urban center at 5:00 AM? Was it during rush hour? Did you include things like construction zones or Halloween costumes or all sorts of things that don’t happen that often, but that you have to handle the right way? If somebody wants to talk about miles, they should also talk about how those miles show they’ve covered the entirety of their operational design domain.

There’s a potential problem with using miles as a measure of progress because they can motivate the wrong behavior. If you’re judged solely on how many miles you’ve racked up, then you’re going to optimize the easy miles perhaps, but worse, every mile on public roads is a chance to make a mistake. Even if you have the best possible safety drivers, every test mile adds some risk. Hopefully that risk is no worse than human driver vehicle risks, but that’s a different discussion. So every mile costs not only money, but also puts you at some risk of some adverse news or some unfortunate event. So you should think carefully about racking up a lot of miles for the sake of miles; you should make every mile earn its keep. Make sure every mile of road testing is actually doing something important.

Now you have to have miles. For sure you won’t know if your system is done until you’ve done some road testing. But the early miles don’t have to be about testing at all, they can be about data collection. You don’t need to run an autonomous system to collect sensor data. You can just collect sensor data with a human driver and have risk no different than someone just driving around normally. That means the road test miles should not be used for primary data collection, but rather as a way to confirm your design is solid. So don’t think of miles on the road is as debugging problems with your system while you’re on public roads. Rather think about road testing miles as a way of making sure you didn’t overlook anything in a system you think is just about ready to go. That means instead of going for road test miles, companies should instead be trying to get data collection miles for most of the miles and making sure they cover the ODD.

Then they can feed that information to a simulation, make sure the system handles all the miles properly, and then after they’re pretty sure they got it right, they should be doing road testing miles simply to confirm that all the engineering effort they put in resulted in a system that behaves the way they expected it to. In other words, road testing miles should be the boring part of tying a ribbon around and putting a bow on a great design. Road testing miles should not be the pointy end of the spear for development.

Going back up to why people talk about miles as a progress metric, sure, having no miles on public roads probably means you’re not ready to deploy because you haven’t checked to make sure it works. But having a ton of miles doesn’t mean you’re ahead, it just means you’re well-funded and you’re out there operating. If you really want to know how someone’s doing, it isn’t just the miles but rather it’s which miles, and it’s how those miles go together with all the other engineering activities to make sure that they’re taking a solid engineering approach to designing a safe self driving car.

For the podcast version of this posting, see: https://archive.org/details/metrics-02-miles-metrics

Thanks to podcast producer Jackie Erickson.

The Lesson Learned from the Tempe Arizona Autonomous Driving System Testing Fatality NTSB Report




Now that the press flurry over the NTSB's report on the Autonomous Driving System (ADS) fatality in Tempe has subsided, it's important to reflect on the lessons to be learned. Hats off to the NTSB for absolutely nailing this. Cheers to the Press who got the messaging right. But not everyone did. The goal of this essay is to help focus on the right lessons to learn, clarify publicly stated misconceptions, and emphasize the most important take-aways.

I encourage everyone in the AV industry to watch the first 5 and a half minutes of the NTSB board meeting video ( Youtube: Link // NTSB Link). Safety leadership should watch the whole thing. Probably twice. Then present a summary at your company's lunch & learn.

Pay particular attention to this part from Chairman Sumwalt: "If your company tests automated driving systems on public roads, this crash -- it was about you.  If you use roads where automated driving systems were being tested, this crash -- it was about you."

I live in Pittsburgh and these public road tests happen near my place of work and my home. I take the lessons from this crash personally. In principle, every time I cross a street I'm potentially placed at risk by any company that might be cutting corners on safety. (I hope that's none. All the companies testing here have voluntarily submitted compliance reports for the well-drafted PennDOT testing guidelines. But not every state has those, and those guidelines were developed largely in response to the fatality we’re discussing.)

I also have long time friends who have invested their careers in this technology. They have brought a vibrant and promising industry to Pittsburgh and other cities.  Negative publicity resulting from a major mishap can threaten the jobs of those employed by those companies.

So it is essential for all of us to get safety right.

The first step: for anyone in charge of testing who doesn't know what a Safety Management System (SMS) is: (A) Watch that NSTB hearing intro. (B) Pause testing on public roads until your company makes a good start down that path. (Again, the PennDOT guidelines are a reasonable first step accepted by a number of companies. LINK)  You’ll sleep better having dramatically improved your company’s safety culture before anyone gets hurt unnecessarily.


Clearing up some misconceptions
I’ve seen some articles and commentary that missed the point of all of this.Large segments of coverage emphasized technical shortcomings of the system - That's not the point. Other coverage highlighted test driver distraction - That's not the point either.  The fatal mishap involved technical shortcomings, and the test driver was not paying adequate attention. Both contributed to the mishap, and both were bad things.

But the lesson to learn is that solid safety culture is without a doubt necessary to prevent avoidable fatalities like these. That is the Point.

To make the most of this teachable moment let's break things down further. These discussions are not really about the particular test platform that was involved. The NTSB report gave that company credit for significant improvement. Rather, the objective is to make sure everyone is focused on ensuring they have learned the most important lesson so we don’t suffer another avoidable ADS testing fatality.

A self-driving car killed someone - NOT THE POINT
This was not a self-driving car. It was a test platform for Automated Driving System (ADS) technology. The difference is night and day.  Any argument that this vehicle was safe to operate on public roads hinged on a human driver not only taking complete responsibility for operational safety, but also being able to intervene when the test vehicle inevitably made a mistake. It's not a fully automated self-driving car if a driver is required to hover with hands above the steering wheel and foot above the brake pedal the entire time the vehicle is operating.

It's a test vehicle. The correct statement is: a test vehicle for developing ADS technology killed someone.

The pedestrian was initially said to jump out of the dark in front of the car - NOT THE POINT
I still hear this sometimes based on the initial video clip that was released immediately after the mishap. The pedestrian walked across almost 4 lanes of road in view of the test vehicle before being struck. The test vehicle detected the pedestrian 5.6 seconds before the crash. That was plenty of time to avoid the crash, and plenty of time to track the pedestrian crossing the street to predict that a crash would occur. Attempting to claim that this crash was unavoidable is incorrect, and won't prevent the next ADS testing fatality.

It's the pedestrian's fault for jaywalking - NOT THE POINT
Jaywalking is what people do when it is 125 yards to the nearest intersection and there is a paved walkway on the median. Even if there is a sign saying not to cross.  Tearing up the paved walkway might help a little on this particular stretch of road, but that's not going to prevent jaywalking as a potential cause of the next ADS testing fatality.

Victim's apparent drug use - NOT THE POINT
It was unlikely that the victim was a fully functional, alert pedestrian. But much of the population isn't in this category for other reasons. Children, distracted walkers, and others with less than perfect capabilities and attention cross the street every day, and we expect drivers to do their best to avoid hitting them.

There is no indication that the victim’s medical condition substantively caused the fatality. (We're back to the fact that the pedestrian did not jump in front of the car.) It would be unreasonable to insist that the public has the responsibility to be fully alert and ready to jump out of the way of an errant ADS test platform at all times they are outside their homes.

Tracking and classification failure - NOT THE POINT
The ADS system on the test vehicle suffered some technical issues that prevented predicting where the pedestrian would be when the test vehicle got there, or even recognizing the object it was sensing was a pedestrian walking a bicycle. However, the point of operating the test vehicle was to find and fix defects.

Defects were expected, and should be expected on other ADS test vehicles. That's why there is a human safety driver. Forbidding public road testing of imperfect ADS systems basically outlaws road testing at this stage. Blaming the technology won't prevent the next ADS testing fatality, but it could hurt the industry for no reason.

It's the technology's fault for ignoring jaywalkers - NOT THE POINT
This idea has been circulating, but apparently this isn't quite true. Jaywalkers aren't ignored, but rather according to the information presented by the NTSB a pedestrian isn't expected to cross the street at first. Once the pedestrian moves for a while a track is built up that could indicate street crossing, but until then movement into the street is considered unexpected if the pedestrian is not at a designated crossing location. A deployment-ready ADS could potentially use a more sophisticated approach to predict when a pedestrian would enter the roadway.

Regardless of implementation, this did not contribute to the fatality because the system never actually classified the victim as a pedestrian. Again, improving this or other ADS technical features won't prevent the next ADS testing fatality. That’s because testing safety is about the safety driver, not which ADS prototype functions happen to be active on any particular test run.

ADS emergency braking behavior - NOT THE POINT
The ADS emergency braking function had behaviors that could hinder its ability to provide backup support to the safety driver. Perhaps another design could have done better for this particular mishap. However, it wasn't the job of the ADS emergency braking to avoid hitting a pedestrian. That was the safety driver's job. Improving ADS emergency braking capabilities might reduce the probability of an ADS testing fatality, but won't entirely prevent the next fatality from happening sooner than it should.

Native emergency braking disabled - NOT THE POINT
It looks bad to have disabled the built-in emergency braking system on the passenger vehicle used as the test platform. The purpose of such systems is to help out after the driver makes a mistake. In this case there is a good, but not perfect, chance it would have helped. But as with the ADS emergency braking function, this simply improves the odds. Any safety expert is going to say your odds are better with both belt and suspenders, but enabling this function alone won't entirely prevent the next ADS testing fatality from happening before it should.

Inattentive safety driver - NOT THE POINT
There is no doubt that an inattentive safety driver is dangerous when supervising an ADS test vehicle. And yet, driver complacency is the expected outcome of asking a human to supervise an automated system that works most of the time. That’s why it’s important to ensure that driver monitoring is done continually and used to provide feedback. (In this case a form of driver monitoring equipment was installed, but data was apparently not used in a way that assured effective driver alertness.)

While enhanced training and stringent driver selection can help, effective analysis and action taken upon monitoring data is required to ensure that drivers are actually paying attention in practice. Simply firing this driver without changing anything else won't prevent the next ADS testing fatality from happening to some other driver who has slipped into bad operational habits.

A fatality is regrettable, but human drivers killed about 100 people that same day with minimal news attention - NOT THE POINT
Some commentators point out the ratio of fatalities caused by test vehicles vs. general automotive fatality rates. They then generally argue that a few deaths in comparison to the ongoing carnage of regular cars is a necessary and appropriate price to pay for progress. However, this argument is not statistically valid.

Consider a reasonable goal that ADS testing (with highly qualified, alert drivers presumed) should be no more dangerous than the risk presented by normal cars. For normal US cars that's ballpark 500 million road miles per pedestrian fatality. This includes mishaps caused by drunk, distracted, and speeding drivers. Due to the far smaller number of miles being driven by current test platform fleet sizes, the "budget" for fatal accidents due to ADS road testing phase should, at this early stage, still be zero.

The fatality somehow “proves” that self-driving car technology isn't viable - NOT THE POINT
Some have tried to draw conclusions about the viability of ADS technology from the fact that there was a testing fatality. However, the issues with ADS technical performance only prove what we already knew. The technology is still maturing, and a human needs to intervene to keep things safe. This crash wasn't about the maturity of the technology; it was about whether the ADS public road testing itself was safe.

Concentrating on technology maturity (for example, via disclosing disengagement rates) serves to focus attention on a long term future of system performance without a safety driver. But the long term isn’t what’s at issue.

The more pressing issue is ensuring that the road testing going on right now is sufficiently safe. At worst, continued use of disengagement rates as the primary metric of ADS performance could hurt safety rather than help. This is because disengagements, if gamed, could incentivize safety drivers to take chances by avoiding disengagements in uncertain situations to make the numbers look better. (Some companies no doubt have strategies to mitigate this risk. But those are probably the companies with an SMS, which is back to the point that matters.)

THE POINT: The safety culture was broken
Safety culture issues were the enabler for this particular crash. Given the limited number of miles that can be accumulated by any current test fleet, we should see no fatalities occur during ADS testing. (Perhaps a truly unavoidable fatality will occur. This is possible, but given the numbers it is unlikely if ADS testing is reasonably safe. So our goal should be set to zero.) Safety culture is critical to ensure this.

The NTSB rightly pushes hard for a safety management system (SMS). But be careful to note that they simply say that this is a part of safety culture, not all of it. Safety culture means, among other things, taking responsibility for ensuring that their safety drivers are actually safe despite the considerable difficulty of accomplishing that. Human safety drivers will make mistakes, but a strong safety culture accounts for such mistakes in ensuring overall safety.

It is important to note that the urgent point here is not regulating self-driving car safety, but rather achieving safe ADS road testing. They are (almost) two entirely different things. Testing safety is about whether the company can consistently put an alert, able-to-react safety driver on the road. On the other hand, ADS safety is about the technology. We need to get to the technology safety part over time, but ADS road testing is the main risk to manage right now.

Perhaps dealing with ADS safety would be easier if the discussions of testing safety and deployment safety were more cleanly separated.

THE TAKEAWAYS:

Chairman Sumwalt summed it up nicely in the intro. (You did watch that 5 and half minute intro, right?)  But to make sure it hits home, this is my take:

One company's crash is every company's crash.  You'll note I didn't name the company involved, because really that's irrelevant to preventing there from being a next fatality and the potential damage it could do to the industry’s reputation.

The bigger point is every company can and should institute good safety culture before further fatalities take place if they have not done so already. The NTSB credited the company at issue with significant change in the right direction.  But it only takes one company who hasn’t gotten the message to be a problem for everyone. We can reasonably expect fatalities involving ADS technology in the future even if these systems are many times safer than human drivers. But there simply aren’t that many vehicles on the road yet for a truly unavoidable mishap to be likely to occur. It’s far too early.

If your company is testing (or plans to test) autonomous vehicles, get a Safety Management System in place before you do public road testing. At least conform to the details in the PennDOT testing guidelines, even if you’re not testing in Pennsylvania. If you are already testing on public roads without an SMS, you should stand down until you get one in place.

Once you have an SMS, consider it a down-payment on a continuing safety culture journey.



Prof. Philip Koopman, Carnegie Mellon University

Author Note: The author and his company work with a variety of customers helping to improve safety. He has been involved with self-driving car safety since the late 1990s. These opinions are his own, and this piece was not sponsored.