SRH Chatbot Evaluation Framework

About the Framework User Guide Evaluation Metrics Chatbot Scorecard Contact Us

Evaluation Metrics

The SRH Chatbot Evaluation Framework assesses chatbots across three high-level dimensions: Product Design, User Experience, and Forward-Looking.

The 33 metrics below provide the basis for assessing chatbot performance. Each metric is assigned a weight of High, Medium, or Low according to its relative importance. You can review the metrics below before completing the Chatbot Scorecard.

01

Product Design

Metrics related to the design, development, privacy, security, and safety of the chatbot.

+

Data Privacy & Security

Obtains informed consent before data collection

High

Informs users about their privacy policy that is clear and using easily accessible language.

High

Communicates data practices transparently, including disclosure of third-party data sharing

High

Adheres to local laws with regard to data collection, storage and privacy

High

Practices data minimization during the onboarding process

Medium

Data collected is only used as necessary for functionality and to improve user experience.

Design

Incorporates co-design and user testing

Medium

For example, pilot testing with adolescent girls and young women before full deployment.

Documents key product and safety decisions and communicates them transparently to users.

Low

Safety

Detects and responds to high-risk, emergency or crisis situations

High

For example, sexual violence or self-harm.

Provides timely referral to appropriate professional services when crisis indicators are identified

High
02

User Experience

Metrics related to accessibility, information quality, performance, and user engagement.

+

Accessibility

Is clear and easy for users to engage with

High

For example, when users are prompted to type a question or choose an option from a menu.

Is affordable and has low barrier to access

Medium

For example, free or low cost and available on widely used platforms.

Accounts for gender and digital inclusion barriers

Medium

Including device sharing, limited connectivity, and different levels of digital literacy.

Is accessible to users with disabilities

Medium

For example, compatibility with screen readers and voice-to-text tools.

Information Reliability and Accuracy

Provides medically accurate information aligned with recognized national or international standards

High

For example, information aligned with WHO guidelines.

Conducts regular content updates aligned with latest SRH guidelines

High

Generates timely responses

High

For example, responses generated in less than one second.

Maintains a low hallucination and/or error rate

High

Performance

Is capable of managing user error

High

The chatbot can manage a query that is not accepted by the chatbot without getting stuck in a loop.

Parent organization has conducted thorough user research to become knowledgeable and conscious of user base and needs

Medium

Purpose of the chatbot is clear from the user perspective

Medium

User Engagement

Uses contextually appropriate tone and language

High

Including culturally relevant terminology and age-appropriate communication.

Allows users to explore topics at their own pace

Medium

Providing opportunities to deepen understanding without encouraging excessive or manipulative engagement.

03

Forward-Looking

Metrics focused on monitoring, learning, outcomes, and responsible de-implementation.

+

Monitoring & Learning

Assessed for safety and accuracy on a regular basis

High

Includes an accessible and easy to use feedback mechanism

High

Allows users to report errors, safety concerns, or malfunction.

Performs periodic analysis of chatbot data and performance indicators

Medium

To identify potential risks and areas for improvement.

Includes a mechanism for individual error spotting and reporting

Medium

For example, thumbs up or down for each individual message to flag oversights.

Continuously reviews advancements in AI

Low

To ensure the chatbot remains safe, medically accurate, and beneficial to users.

Short-Term Outcomes

Offers in person / real life resources and services for users where necessary

High

Increases users' SRH knowledge

Medium

Tracks clicks through to in person / real life resources and services

Medium

For example, tracking intent to seek care.

Provides actionable tools to increase users' digital literacy and confidence online

Low

For example, supporting users' AI literacy and teaching users to use AI tools responsibly.

Termination & De-Implementation

Includes contingency plan for continuity of care

Medium

Includes a plan for where to send the audience if the chatbot shuts down.

Has strong plan for data handling in the event of a shutdown

Medium

Includes what happens to stored data that is being used to keep users safe.