SRH Chatbot Evaluation Framework

About the Framework User Guide Evaluation Metrics Chatbot Scorecard Contact Us

About the Framework

This project was a collaboration between graduate students at the School of International and Public Affairs (SIPA), Columbia University, under the guidance of Professor Savita Bailur, and Girl Effect to develop a practical framework for assessing the quality of chatbots designed to deliver sexual and reproductive health (SRH) information to girls and young women in the Global South. While AI-enabled SRH chatbots are rapidly expanding, there is a lack of standardized mechanisms for assessing whether these tools are safe, effective, and responsive to users’ needs.

To address this gap, we undertook the following:

  • An extensive literature review
  • Key informant interviews with eleven experts across the fields of artificial intelligence, sexual and reproductive health, digital design, and gender
  • Designed, built and tested the metrics and scorecard for evaluating an SRH chatbot

These findings informed a scorecard spanning 30+ weighted metrics across areas including:

  • Product design
  • Information reliability
  • Privacy and security
  • User safety
  • Accessibility
  • Monitoring and learning
  • Sustainability

The framework is intended as a public good, freely available to help strengthen standards, accountability, and quality across the growing field of SRH chatbots. Our online scorecard provides these stakeholders with a shared benchmark for assessing and strengthening SRH chatbots, helping establish a more consistent standard for quality across the field.

However, we also acknowledge that such metrics can be reductivist and subjective. First, our goal is that such a scorecard can prompt the evaluator’s thinking (and to that end, we have also added a free text box to capture reflections at the end of each of the three buckets of Product Design, User Experience and Forward-thinking). Second, stay tuned for a series of blogs around this exercise, which discuss issues from consent to de-implementation.

Who is this framework for?

The framework is designed for funders, chatbot developers, implementers, and other stakeholders involved in the design, financing, deployment, and evaluation of SRH chatbots.

For funders, it can provide a common benchmark for assessing whether proposed or existing chatbot initiatives meet essential standards for quality, safety, privacy, and accountability, and can inform investment decisions. For developers and implementers, it can support human-centered design from the outset by encouraging user research, co-design, user testing, feedback mechanisms, and meaningful connections to in-person services.

Acknowledgments

We would like to extend our sincere thanks to our co-collaborators at Girl Effect (Alex Fulcher and Soma Mitra-Behura), and the following experts for their time and insights: Abhijit Mali, Ben Burrows, Daniel Bjorkegren, Divya Panchaksharappa Budihal, Isabelle Amazon-Brown, Kandyce Brennan, Louisa Rosenheck, Sara Chamberlain, Sofie Meyer, Tze Wei Liew, and Zezhen Wu.