NIST: Identified risks are assessed, analyzed, or tracked.
MEASURE 1: Appropriate methods and metrics are identified and applied.
MEASURE 1 has 3 subcategories. MEASURE 1 in plain English, with example evidence
MEASURE 1.1
Approaches and metrics for measurement of AI risks enumerated during the MAP function are selected for implementation starting with the most significant AI risks. The risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented.
In plain English: Pick metrics for the most significant mapped risks first, and record what you can't measure.
First 2 of 12 suggested actions in the Playbook:
- Establish approaches for detecting, tracking and measuring known risks, errors, incidents or negative impacts.
- Identify testing procedures and metrics to demonstrate whether or not the system is fit for purpose and functioning as claimed.
Tagged by NIST for: AI Development, TEVV, Domain Experts
MEASURE 1.2
Appropriateness of AI metrics and effectiveness of existing controls are regularly assessed and updated, including reports of errors and potential impacts on affected communities.
In plain English: Check regularly that your metrics and controls still work, including error reports and effects on communities.
First 2 of 8 suggested actions in the Playbook:
- Assess external validity of all measurements (e.g., the degree to which measurements taken in one context can generalize to other contexts).
- Assess effectiveness of existing metrics and controls on a regular basis throughout the AI system lifecycle.
Tagged by NIST for: TEVV, AI Impact Assessment, AI Development, AI Deployment, Affected Individuals and Communities
MEASURE 1.3
Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates. Domain experts, users, AI actors external to the team that developed or deployed the AI system, and affected communities are consulted in support of assessments as necessary per organizational risk tolerance.
In plain English: Have assessments done by people who didn't build the system, consulting outside experts and affected groups where needed.
First 2 of 8 suggested actions in the Playbook:
- Evaluate TEVV processes regarding incentives to identify risks and impacts.
- Utilize separate testing teams established in the Govern function (2.1 and 4.1) to enable independent decisions and course-correction for AI systems. Track processes and measure and document change in performance.
Tagged by NIST for: TEVV, AI Impact Assessment, AI Development, AI Deployment, Affected Individuals and Communities, Domain Experts, End-Users, Operation and Monitoring
MEASURE 2: AI systems are evaluated for trustworthy characteristics.
MEASURE 2 has 13 subcategories. MEASURE 2 in plain English, with example evidence
MEASURE 2.1
Test sets, metrics, and details about the tools used during TEVV are documented.
In plain English: Document the test sets, metrics and tools used in testing.
First 2 of 3 suggested actions in the Playbook:
- Leverage existing industry best practices for transparency and documentation of all possible aspects of measurements. Examples include: data sheet for data sets, model cards
- Regularly assess the effectiveness of tools used to document measurement approaches, test sets, metrics, processes and materials used
Tagged by NIST for: TEVV
MEASURE 2.2
Evaluations involving human subjects meet applicable requirements (including human subject protection) and are representative of the relevant population.
In plain English: Make tests with human participants meet protection rules and represent the real population.
First 2 of 8 suggested actions in the Playbook:
- Follow human subjects research requirements as established by organizational and disciplinary requirements, including informed consent and compensation, during dataset collection activities.
- Analyze differences between intended and actual population of users or data subjects, including likelihood for errors, incidents or negative impacts.
Tagged by NIST for: TEVV, Human Factors, AI Development
MEASURE 2.3
AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s). Measures are documented.
In plain English: Measure performance in conditions like the real deployment, and record the results.
First 2 of 9 suggested actions in the Playbook:
- Conduct regular and sustained engagement with potentially impacted communities
- Maintain a demographically diverse and multidisciplinary and collaborative internal team
Tagged by NIST for: TEVV, AI Deployment
MEASURE 2.4
The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production.
In plain English: Monitor how the system and its components behave in production.
First 2 of 7 suggested actions in the Playbook:
- Monitor and document how metrics and performance indicators observed in production differ from the same metrics collected during pre-deployment testing. When differences are observed, consider error propagation and feedback loop risks.
- Utilize hypothesis testing or human domain expertise to measure monitored distribution differences in new input or output data relative to test environments
Tagged by NIST for: AI Deployment, TEVV
MEASURE 2.5
The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented.
In plain English: Show the system is valid and reliable, and record where it may not generalise.
First 2 of 15 suggested actions in the Playbook:
- Define the operating conditions and socio-technical context under which the AI system will be validated.
- Define and document processes to establish the system’s operational conditions and limits.
Tagged by NIST for: TEVV, Domain Experts
MEASURE 2.6
The AI system is evaluated regularly for safety risks – as identified in the MAP function. The AI system to be deployed is demonstrated to be safe, its residual negative risk does not exceed the risk tolerance, and it can fail safely, particularly if made to operate beyond its knowledge limits. Safety metrics reflect system reliability and robustness, real-time monitoring, and response times for AI system failures.
In plain English: Test safety regularly: residual risk within tolerance, and the system fails safely outside its limits.
First 2 of 7 suggested actions in the Playbook:
- Thoroughly measure system performance in development and deployment contexts, and under stress conditions.
- Employ test data assessments and simulations before proceeding to production testing. Track multiple performance quality and error metrics.
- Stress-test system performance under likely scenarios (e.g., concept drift, high load) and beyond known limitations, in consultation with domain experts.
- Test the system under conditions similar to those related to past known incidents or near-misses and measure system performance and safety characteristics
- Apply chaos engineering approaches to test systems in extreme conditions and gauge unexpected responses.
- Document the range of conditions under which the system has been tested and demonstrated to fail safely.
- Measure and monitor system performance in real-time to enable rapid response when AI system incidents are detected.
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment
MEASURE 2.7
AI system security and resilience – as identified in the MAP function – are evaluated and documented.
In plain English: Evaluate and record the system's security and resilience.
First 2 of 10 suggested actions in the Playbook:
- Establish and track AI system security tests and metrics (e.g., red-teaming activities, frequency and rate of anomalous events, system down-time, incident response times, time-to-bypass, etc.).
- Use red-team exercises to actively test the system under adversarial or stress conditions, measure system response, assess failure modes or determine if system can return to normal function after an unexpected adverse event.
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment
MEASURE 2.8
Risks associated with transparency and accountability – as identified in the MAP function – are examined and documented.
In plain English: Examine and record transparency and accountability risks.
First 2 of 6 suggested actions in the Playbook:
- Instrument the system for measurement and tracking, e.g., by maintaining histories, audit logs and other information that can be used by AI actors to review and evaluate possible sources of error, bias, or vulnerability.
- Calibrate controls for users in close collaboration with experts in user interaction and user experience (UI/UX), human computer interaction (HCI), and/or human-AI teaming.
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment
MEASURE 2.9
The AI model is explained, validated, and documented, and AI system output is interpreted within its context – as identified in the MAP function – to inform responsible use and governance.
In plain English: Explain and validate the model, and interpret its output in context.
First 2 of 11 suggested actions in the Playbook:
- Verify systems are developed to produce explainable models, post-hoc explanations and audit logs.
- When possible or available, utilize approaches that are inherently explainable, such as traditional and penalized generalized linear models , decision trees, nearest-neighbor and prototype-based approaches, rule-based models, generalized additive models , explainable boosting machines and neural additive models.
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment, End-Users
MEASURE 2.10
Privacy risk of the AI system – as identified in the MAP function – is examined and documented.
In plain English: Examine and record the system's privacy risk.
First 2 of 8 suggested actions in the Playbook:
- Specify privacy-related values, frameworks, and attributes that are applicable in the context of use through direct engagement with end users and potentially impacted groups and communities.
- Document collection, use, management, and disclosure of personally sensitive information in datasets, in accordance with privacy and data governance policies
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment, End-Users
MEASURE 2.11
Fairness and bias – as identified in the MAP function – are evaluated and results are documented.
In plain English: Evaluate fairness and bias, and record the results.
First 2 of 24 suggested actions in the Playbook:
- Conduct fairness assessments to manage computational and statistical forms of bias which include the following steps:
- Identify types of harms, including allocational, representational, quality of service, stereotyping, or erasure
- Identify across, within, and intersecting groups that might be harmed
- Quantify harms using both a general fairness metric, if appropriate (e.g. demographic parity, equalized odds, equal opportunity, statistical hypothesis tests), and custom, context-specific metrics developed in collaboration with affected communities
- Analyze quantified harms for contextually significant differences across groups, within groups, and among intersecting groups
- Refine identification of within-group and intersectional group disparities.
- Evaluate underlying data distributions and employ sensitivity analysis during the analysis of quantified harms.
- Evaluate quality metrics including false positive rates and false negative rates.
- Consider biases affecting small groups, within-group or intersectional communities, or single individuals.
- Understand and consider sources of bias in training and TEVV data:
- Differences in distributions of outcomes across and within groups, including intersecting groups.
- Completeness, representativeness and balance of data sources.
- Identify input data features that may serve as proxies for demographic group membership (i.e., credit score, ZIP code) or otherwise give rise to emergent bias within AI systems.
- Forms of systemic bias in images, text (or word embeddings), audio or other complex or unstructured data.
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment, End-Users, Affected Individuals and Communities
MEASURE 2.12
Environmental impact and sustainability of AI model training and management activities – as identified in the MAP function – are assessed and documented.
In plain English: Assess and record the environmental impact of training and running the models.
First 2 of 6 suggested actions in the Playbook:
- Include environmental impact indicators in AI system design and development plans, including reducing consumption and improving efficiencies.
- Identify and implement key indicators of AI system energy and water consumption and efficiency, and/or GHG emissions.
Tagged by NIST for: TEVV, Domain Experts, Operation and Monitoring, AI Impact Assessment, AI Deployment
MEASURE 2.13
Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented.
In plain English: Check whether your testing metrics and processes actually work.
First 2 of 4 suggested actions in the Playbook:
- Review selected system metrics and associated TEVV processes to determine if they are able to sustain system improvements, including the identification and removal of errors.
- Regularly evaluate system metrics for utility, and consider descriptive approaches in place of overly complex methods.
Tagged by NIST for: TEVV, AI Deployment, Operation and Monitoring
MEASURE 3: Mechanisms for tracking identified AI risks over time are in place.
MEASURE 3 has 3 subcategories. MEASURE 3 in plain English, with example evidence
MEASURE 3.1
Approaches, personnel, and documentation are in place to regularly identify and track existing, unanticipated, and emergent AI risks based on factors such as intended and actual performance in deployed contexts.
In plain English: Track existing, unexpected and new risks once the system is live.
First 2 of 6 suggested actions in the Playbook:
- Compare AI system risks with:
- simpler or traditional models
- human baseline performance
- other manual performance benchmarks
- Compare end user and community feedback about deployed AI systems to internal measures of system performance.
Tagged by NIST for: TEVV, AI Impact Assessment, Operation and Monitoring
MEASURE 3.2
Risk tracking approaches are considered for settings where AI risks are difficult to assess using currently available measurement techniques or where metrics are not yet available.
In plain English: Have a way to track risks you can't yet measure well.
First 2 of 3 suggested actions in the Playbook:
- Establish processes for tracking emergent risks that may not be measurable with current approaches. Some processes may include:
- Recourse mechanisms for faulty AI system outputs.
- Bug bounties.
- Human-centered design approaches.
- User-interaction and experience research.
- Participatory stakeholder engagement with affected or potentially impacted individuals and communities.
- Identify AI actors responsible for tracking emergent risks and inventory methods.
Tagged by NIST for: TEVV, Domain Experts, AI Impact Assessment, Operation and Monitoring
MEASURE 3.3
Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics.
In plain English: Let users and affected people report problems and appeal outcomes, and count those reports in evaluation.
First 2 of 6 suggested actions in the Playbook:
- Measure efficacy of end user and operator error reporting processes.
- Categorize and analyze type and rate of end user appeal requests and results.
Tagged by NIST for: TEVV, AI Deployment, Operation and Monitoring, End-Users, Affected Individuals and Communities
MEASURE 4: Feedback about efficacy of measurement is gathered and assessed.
MEASURE 4 has 3 subcategories. MEASURE 4 in plain English, with example evidence
MEASURE 4.1
Measurement approaches for identifying AI risks are connected to deployment context(s) and informed through consultation with domain experts and other end users. Approaches are documented.
In plain English: Tie measurement to the real deployment context, with input from domain experts and end users.
First 2 of 6 suggested actions in the Playbook:
- Support mechanisms for capturing feedback from system end users (including domain experts, operators, and practitioners). Successful approaches are:
- conducted in settings where end users are able to openly share their doubts and insights about AI system output, and in connection to their specific context of use (including setting and task-specific lines of inquiry)
- developed and implemented by human-factors and socio-technical domain experts and researchers
- designed to ensure control of interviewer and end user subjectivity and biases
- Identify and document approaches
- for evaluating and integrating elicited feedback from system end users
- in collaboration with human-factors and socio-technical domain experts,
- to actively inform a process of continual improvement.
Tagged by NIST for: TEVV, AI Deployment, Operation and Monitoring, End-Users, Affected Individuals and Communities
MEASURE 4.2
Measurement results regarding AI system trustworthiness in deployment context(s) and across the AI lifecycle are informed by input from domain experts and relevant AI actors to validate whether the system is performing consistently as intended. Results are documented.
In plain English: Check with experts and the people involved that the system performs as intended once deployed.
First 2 of 6 suggested actions in the Playbook:
- Integrate feedback from end users, operators, and affected individuals and communities from Map function as inputs to assess AI system trustworthiness characteristics. Ensure both positive and negative feedback is being assessed.
- Evaluate feedback in connection with AI system trustworthiness characteristics from Measure 2.5 to 2.11.
Tagged by NIST for: TEVV, AI Deployment, Domain Experts, Operation and Monitoring, End-Users
MEASURE 4.3
Measurable performance improvements or declines based on consultations with relevant AI actors, including affected communities, and field data about context-relevant risks and trustworthiness characteristics are identified and documented.
In plain English: Record measurable improvements or declines, from consultation and field data.
First 2 of 6 suggested actions in the Playbook:
- Develop baseline quantitative measures for trustworthy characteristics.
- Delimit and characterize baseline operation values and states.
Tagged by NIST for: TEVV, AI Deployment, Operation and Monitoring, End-Users, Affected Individuals and Communities