SRH Chatbot Evaluation Framework
Evaluation Metrics
The SRH Chatbot Evaluation Framework assesses chatbots across three high-level dimensions: Product Design, User Experience, and Forward-Looking.
The 33 metrics below provide the basis for assessing chatbot performance. Each metric is assigned a weight of High, Medium, or Low according to its relative importance. You can review the metrics below before completing the Chatbot Scorecard.
01
Product Design
Metrics related to the design, development, privacy, security,
and safety of the chatbot.
+
Product Design
Metrics related to the design, development, privacy, security, and safety of the chatbot.
Data Privacy & Security
Informs users about their privacy policy that is clear and using easily accessible language.
HighCommunicates data practices transparently, including disclosure of third-party data sharing
HighAdheres to local laws with regard to data collection, storage and privacy
HighPractices data minimization during the onboarding process
MediumData collected is only used as necessary for functionality and to improve user experience.
Design
Incorporates co-design and user testing
MediumFor example, pilot testing with adolescent girls and young women before full deployment.
Documents key product and safety decisions and communicates them transparently to users.
LowSafety
Detects and responds to high-risk, emergency or crisis situations
HighFor example, sexual violence or self-harm.
Provides timely referral to appropriate professional services when crisis indicators are identified
High
02
User Experience
Metrics related to accessibility, information quality,
performance, and user engagement.
+
User Experience
Metrics related to accessibility, information quality, performance, and user engagement.
Accessibility
Is clear and easy for users to engage with
HighFor example, when users are prompted to type a question or choose an option from a menu.
Is affordable and has low barrier to access
MediumFor example, free or low cost and available on widely used platforms.
Accounts for gender and digital inclusion barriers
MediumIncluding device sharing, limited connectivity, and different levels of digital literacy.
Is accessible to users with disabilities
MediumFor example, compatibility with screen readers and voice-to-text tools.
Information Reliability and Accuracy
Provides medically accurate information aligned with recognized national or international standards
HighFor example, information aligned with WHO guidelines.
Conducts regular content updates aligned with latest SRH guidelines
HighGenerates timely responses
HighFor example, responses generated in less than one second.
Maintains a low hallucination and/or error rate
HighPerformance
Is capable of managing user error
HighThe chatbot can manage a query that is not accepted by the chatbot without getting stuck in a loop.
Parent organization has conducted thorough user research to become knowledgeable and conscious of user base and needs
MediumPurpose of the chatbot is clear from the user perspective
MediumUser Engagement
Uses contextually appropriate tone and language
HighIncluding culturally relevant terminology and age-appropriate communication.
Allows users to explore topics at their own pace
MediumProviding opportunities to deepen understanding without encouraging excessive or manipulative engagement.
03
Forward-Looking
Metrics focused on monitoring, learning, outcomes, and
responsible de-implementation.
+
Forward-Looking
Metrics focused on monitoring, learning, outcomes, and responsible de-implementation.
Monitoring & Learning
Assessed for safety and accuracy on a regular basis
HighIncludes an accessible and easy to use feedback mechanism
HighAllows users to report errors, safety concerns, or malfunction.
Performs periodic analysis of chatbot data and performance indicators
MediumTo identify potential risks and areas for improvement.
Includes a mechanism for individual error spotting and reporting
MediumFor example, thumbs up or down for each individual message to flag oversights.
Continuously reviews advancements in AI
LowTo ensure the chatbot remains safe, medically accurate, and beneficial to users.
Short-Term Outcomes
Offers in person / real life resources and services for users where necessary
HighIncreases users' SRH knowledge
MediumTracks clicks through to in person / real life resources and services
MediumFor example, tracking intent to seek care.
Provides actionable tools to increase users' digital literacy and confidence online
LowFor example, supporting users' AI literacy and teaching users to use AI tools responsibly.
Termination & De-Implementation
Includes contingency plan for continuity of care
MediumIncludes a plan for where to send the audience if the chatbot shuts down.
Has strong plan for data handling in the event of a shutdown
MediumIncludes what happens to stored data that is being used to keep users safe.