Beyond Benchmarks: Lessons from Participatory AI Evaluation at Deep Learning Indaba 2026

Lewinsky
Lewinsky
09 September 2026

Facilitating the Beyond Benchmarks workshop at Deep Learning Indaba 2026.

Beyond Benchmarks: Lessons from Participatory AI Evaluation at Deep Learning Indaba 2026

Exploring what community-driven, multi-turn evaluation can reveal about AI systems in African contexts 

AI evaluation is still largely dominated by benchmarks: standardized datasets and metrics designed to provide comparable measures of model performance.

Benchmarks remain important, but they can tell us only part of the story.

They are often poorly equipped to reveal what happens when people interact with AI systems over multiple conversational turns, communicate in different languages or dialects, introduce real-world constraints, or judge responses according to what is useful and appropriate in their own context.

At Deep Learning Indaba 2026, YUX Cultural AI Lab, Kitala AI, and Microsoft Research Africa hosted Beyond Benchmarks: Scaling Multi-Turn Participatory AI Evaluations in Health and Education, a 90-minute hands-on workshop exploring an alternative and complementary approach.

Facilitating the Beyond Benchmarks workshop at Deep Learning Indaba 2026.

Facilitating the Beyond Benchmarks workshop at Deep Learning Indaba 2026.

Participants worked in small groups to construct realistic scenarios, introduce adversarial and contextual challenges, interact with AI systems through Kitala.ai, and collectively evaluate the resulting conversations.

The workshop built on participatory evaluation work in Rwanda, Kenya, and Senegal.

Participants during the Beyond Benchmarks workshop at Deep Learning Indaba 2026.

Participants during the Beyond Benchmarks workshop at Deep Learning Indaba 2026.

Why go beyond benchmarks?

Benchmarks are foundational to AI evaluation. They allow researchers to systematically assess capabilities such as accuracy, robustness, safety, and toxicity, compare performance across models, and track improvements over time.

However, these measures largely assess performance under predefined and controlled conditions.

They tell us comparatively little about what happens when AI systems encounter the complexity of real-world human interaction.

This distinction is increasingly important as AI moves from being evaluated primarily as a model to being experienced as a product embedded within people's everyday lives.

Strong model performance does not necessarily translate into a successful user interaction, and a successful interaction does not necessarily translate into meaningful real-world impact.

Human-in-the-Loop approaches respond to this gap by combining technical evaluation with human judgement, allowing AI behaviour to be assessed in relation to user expectations, experiences, and ethical considerations.

This becomes especially consequential in African contexts.

AI systems trained predominantly on Global North datasets may struggle with local communication practices, cultural expressions, healthcare realities, and other forms of contextual knowledge.

Across linguistically diverse environments, people may also move fluidly between languages, use culturally specific expressions, or communicate in ways that differ substantially from the standardized prompts on which models are typically evaluated.

Real-world AI interactions are also rarely benchmark-shaped.

A person may begin with an ambiguous question, provide important context only after receiving a response, challenge the system's recommendation, introduce a financial or social constraint, switch languages midway through the conversation, or change what they are asking based on what the AI previously said.

These dynamics are particularly difficult to capture through static, single-turn evaluation.

So the question is not whether participatory evaluation should replace benchmarks.

Benchmarks remain essential for standardized and scalable measurement.

Rather, the challenge is understanding what benchmarks cannot see, and what additional forms of evidence are needed to understand AI performance in context.

This leads to the broader question that guided our workshop:

How do we evaluate not only whether an AI system can produce the right answer, but whether it can support the right interaction for a particular person, situation, language, and context?

Putting multi-turn evaluation into practice

The workshop brought together participants with backgrounds including AI and ML engineering, AI safety, academia, ethics, consulting, software engineering, and students and recent graduates.

Participants represented or had connections to several African countries, including Nigeria, Kenya, Ghana, Rwanda, Senegal, Burkina Faso, South Africa, and the Democratic Republic of Congo, alongside participants from outside the continent.

This diversity became especially important during the evaluation activity because participants were able to test AI interactions across multiple languages, including Yoruba, Igbo, Twi, isiZulu, Swahili, French, Arabic, Amharic and others.

A central feature of the workshop was multi-turn evaluation.

Participants formed small groups and selected scenario cards representing situations in health or education. They also selected contextual or adversarial cards designed to introduce complications into the interaction.

Rather than submitting unrelated prompts, groups constructed conversations in which each question could build on what had happened previously.

 

Scenario Cards.

Scenario Cards.

 

Participants then introduced additional complications through adversarial nudges.

Adversarial Nudges.

Adversarial Nudges.

Groups collectively evaluated the resulting interactions through Kitala.ai using conversational metrics and response-level flags.

This shifted evaluation from a static prompt → response → score structure toward evaluation of the interaction as a whole.

What did we find?

Language quickly became a major evaluation surface

One of the clearest observations from the workshop was how strongly participants gravitated toward testing language capabilities.

Participants experimented with several African languages and language varieties, revealing issues that would be difficult to understand from an English-language benchmark alone.

In Yoruba interactions, facilitators observed responses that could be vague, unnecessarily long, or indirect. Participants also encountered differences associated with Yoruba dialects.

An Igbo interaction produced an especially visible failure: the system did not appear to understand the language and generated text that participants perceived as resembling Chinese.

Another group reported difficulty with the Arabic dialect they selected.

Participants also explored code-switching. In one case, the model did not naturally code-switch between Swahili and English until explicitly instructed to do so.

These examples illustrate an important distinction between nominal language support and interactional language competence.

A system may technically support a language while still struggling with dialects, code-switching, follow-up questions, conversational naturalness, or the way speakers actually use that language.

This suggests that multilingual evaluation needs to ask more than:

Does the model speak this language?

It must also ask:

How well can the model sustain a useful conversation in the ways people actually communicate?

 

Participants testing AI systems during the workshop

Participants testing AI systems during the workshop

Multi-turn evaluation exposed different failures

The workshop also demonstrated why evaluating conversations rather than isolated responses matters.

In one Yoruba interaction involving an adversarial statement around job changes, facilitators observed that the model's behaviour deteriorated as the conversation progressed.

Responses became longer and the model appeared to reinforce problematic behaviour rather than redirecting the interaction appropriately.

Whether such behaviour represents a systematic model failure would require considerably more controlled testing.

However, the interaction demonstrates the diagnostic value of multi-turn evaluation: some problems become visible only after the user challenges the system, adds context, or moves the conversation in an unexpected direction.

This is particularly relevant for deployed conversational systems.

Real users do not necessarily communicate through perfectly constructed prompts. They contradict themselves, challenge recommendations, introduce new information, express frustration, switch languages and test boundaries.

Evaluation methods need to accommodate that messiness.

Participatory evaluation can reveal what researchers did not think to test

Participants generally responded positively to the idea of moving beyond benchmark-only evaluation.

One participant described the approach as useful because it brings researchers closer to the end user.

In conventional evaluation, researchers often determine in advance what should be tested, which prompts represent the task, what constitutes a failure, which metrics matter, and what successful performance looks like.

Participatory evaluation redistributes some of that evaluative authority.

Instead of researchers being the only people deciding what constitutes an important edge case, people with relevant linguistic, cultural, professional, or lived knowledge can actively challenge the system.

This doesn't make participatory evaluation inherently unbiased or representative.

Rather, it can expand the range of things an evaluation is capable of seeing.

But participatory evaluation has its own biases

One of the most valuable workshop findings was actually a critique of the evaluation method itself.

Kitala included predefined flags participants could apply to model responses. Many of these flags focused on failures.

Participants questioned this directly:

Why were the flags biased toward negative outcomes?

If an evaluation interface primarily presents evaluators with categories such as inaccurate, unsafe, irrelevant, or confusing, are we inadvertently instructing them to search for those failures?

The evaluation instrument itself can shape what evaluators notice.

Participants therefore suggested that evaluations should also make space for positive signals: moments where the model understood context particularly well, adapted appropriately, communicated clearly, provided a strong follow-up, or handled a difficult situation effectively.

This changes the question from:

What went wrong?

to:

What happened in this interaction that is important enough for us to understand?

That includes both successes and failures.

Evaluation burden matters too

Another challenge was the conversational feedback loop.

Having participants interact with the model, evaluate responses, provide comments, continue the conversation, and repeat this process can become time-consuming.

This raises a design tension between evaluation richness and interaction naturalness.

The more questions researchers insert into an AI interaction, the less the resulting conversation resembles how someone might naturally use that AI system.

A structured evaluation may require participants to deliberately test particular scenarios and rate responses. A naturalistic evaluation may instead allow people to use AI freely and collect much lighter feedback.

Neither is inherently better.

They answer different questions.

Beyond benchmarks should not mean without benchmarks

One of the strongest conclusions from the workshop is that "beyond benchmarks" should not mean without benchmarks.

Instead, different evaluation approaches answer different questions.

Evaluation Layer

What it helps us understand

Benchmarks & automated evaluation

Can the system perform a defined task consistently at scale?

Expert evaluation

Is the response technically/domain appropriate, accurate and safe?

Participatory evaluation

Does the interaction work for people within their language, circumstances and lived context?

Naturalistic evaluation

When people are free to use the system themselves, when and how does it actually enter their lives?

 

The opportunity is not to select one layer but rather to understand where each layer is strong, what each one misses, and how evidence can be triangulated across them.

This is particularly important when evaluating AI systems intended for African contexts, where linguistic diversity, infrastructure, access, cultural context and local realities can expose limitations invisible within standardized evaluation environments.

What does better participatory evaluation look like?

The workshop points toward several practical principles for future work:

  • Balance strengths and failures. Evaluation interfaces should allow participants to identify strong responses as readily as problematic ones.
  • Keep structured metrics small and interpretable. Too many overlapping metrics increase evaluator burden and reduce consistency.
  • Leave room for explanation. Flags tell us what category something belongs to; qualitative reflections tell us why it matters.
  • Match feedback modality to interaction modality. Voice-based interactions may benefit from voice-based reflections rather than requiring lengthy written feedback.
  • Combine evaluator perspectives. Community members, domain experts and researchers bring different forms of expertise.
  • Design for natural behaviour where appropriate. Highly instrumented evaluation is useful for controlled testing, but lighter-touch longitudinal methods may be needed to understand real-world use.

From evaluation platform to evaluation infrastructure

Participants also pushed the discussion beyond the workshop methodology toward the role of evaluation platforms themselves.

One participant described Kitala as a kind of "white paint" or base layer that could potentially sit underneath AI products rather than existing only as a separate evaluation environment.

Participants imagined a more integrated evaluation lifecycle in which researchers could: design an evaluation → recruit or engage evaluators → collect interactions and feedback → analyse patterns → compare systems within one environment. 

Others questioned whether participatory feedback could eventually operate more quietly in the background of everyday AI interactions rather than requiring users to enter a dedicated evaluation platform.

These ideas point toward a broader design opportunity:

What would it look like if participatory evaluation were treated not as a one-time research activity, but as ongoing infrastructure for understanding deployed AI systems?

Kitala AI interface

Kitala AI interface

Beyond a single score

The workshop started from a relatively straightforward proposition: benchmarks cannot tell us everything we need to know about AI systems.

The hands-on evaluation made the implications of that proposition much more concrete.

Participants surfaced multilingual failures, contextual shortcomings, conversational issues and unexpected model behaviours.

But they also identified limitations in the participatory evaluation approach itself—from negatively biased flags and restrictive taxonomies to feedback burden and the difficulty of maintaining natural conversation while simultaneously evaluating it.

Going beyond benchmarks does not simply mean adding humans to AI evaluation.

It requires thinking carefully about which humans participate, what knowledge they bring, what situations they test, what evaluative categories researchers give them, how feedback is collected, and whose definition of successful AI ultimately counts.

For AI systems deployed across diverse African contexts, evaluation therefore needs to move beyond asking only:

"How well does the model perform?"

toward asking:

"For whom does it perform well, under what circumstances, in what language, according to whose judgement; and what does our evaluation method itself prevent us from seeing?"

Participatory evaluation offers one route toward answering those questions.

The challenge now is to develop approaches that are rigorous enough to generate comparable evidence while remaining open enough for communities to surface the things researchers did not know to measure in the first place.

The Beyond Benchmarks workshop was one step toward that goal.