Data

How capable is AI in language education? Explore our open benchmark dataset

James Edgell Isaac Pattis Ben Knight James Edgell, Isaac Pattis & Ben Knight
23 Jul 2026 5 min read
Share:

How capable are different LLMs at helping someone learn a language?

The L2-Bench project was set up to answer this question, and now you can find out for yourself.

Last week we shared how nine large language models (LLMs) performed across 1,000 real language-teaching tasks constituting our L2-Bench benchmark. This week we’re releasing the open dataset behind those results, free for anyone to explore.

So what exactly are we sharing? L2-Bench is one of the first open benchmarks for evaluating AI capabilities in second language education, and one of the most extensively validated. At its heart are:

  • 1,000 real-world second language education tasks – the kinds of things practitioners are involved with every day, from planning a lesson to giving a learner feedback
  • Task tagging to a 12-competency framework (with 31 sub-competencies) that describe the range of roles that intentionally design the conditions that shape how people learn languages effectively
  • An expert scoring rubric of pass/fail criteria for every task, consisting of task-specific, consensus, and universal criteria
  • A “gold-standard” reference answer showing what a strong single-turn response looks like for each task.

Depending on your interests, you could use the L2-Bench dataset to:

  • See how well any AI system handles the language learning tasks you care about most
  • Explore where models do well — and where they still struggle — for different learners, roles, geographies, and settings
  • Replicate the L2-Bench evaluation process to get a score for any LLM
  • Re-purpose the L2-Bench task and rubric design components to build AI evaluations in your domain.

Releasing our AI evaluation benchmark dataset for language education

Benchmarks do more than measure AI. They shape the future of AI by influencing what technology companies optimise for, what researchers prioritise, and ultimately what capabilities become embedded within products.

AI is already part of the language learning tools used by millions of learners worldwide, and we think the best way to make those tools safer and more effective is to work collaboratively and openly with other educational practitioners, institutions, and researchers.

Therefore, we are releasing L2-Bench as an open dataset, validated by more than 200 experienced educators from 45 countries, to support an AI for Education evaluation ecosystem.

How can I access the open L2-Bench dataset?

The L2-Bench benchmark dataset is available now on HuggingFace:

  • See the dataset card for further information.
  • We have also open-sourced our scoring pipeline, allowing anyone to run models and produce evaluations on L2-Bench.

Want the background before you dive in? Our two earlier posts explain how L2-Bench came together, including links to accompanying papers on methods (how we built L2-Bench) and results (how LLMs scored on L2-Bench):

If you’d like an alternative distribution, please get in touch via the “Register Interest” form below.

What kinds of tasks are in L2-Bench?

In its 1,000 tasks, L2-Bench attempts to address the challenges of representing the complexity of language education around the world.

The tasks are spread fairly evenly across all 12 teaching competencies — roughly 80 to 110 tasks each — covering everything from course and lesson planning to giving feedback, running activities, and assessing progress.

Crucially, with each task tagged with context factors that contribute most to real learning environments, you can filter to the situations that matter to you, including:

  • The age and proficiency level of learners
  • The learner’s first language
  • The teaching setting and resources available.

The tasks were designed to reflect classrooms across 90+ countries, which means you can explore how AI performs for very different learners and settings — not just one idealised classroom.

World map showing the geographic distribution of L2-Bench tasks; darker shading indicates countries with more tasks set in that location.
Geographic distribution of L2-Bench tasks. Tasks were sampled according to a weighting that accounted for countries learning English, population size, and English language proficiency.

However, it’s important to be aware of limitations in this initial L2-Bench dataset that we will address in future work. Please read on below for how to get involved:

  • Since L2-Bench items are “single-turn”, we cannot directly observe behaviours that only emerge over a dialogue, meaning some scores may under- or over-estimate a model’s capability
  • L2-Bench covers English as the target language and is grounded in pedagogy frameworks of European origin, which may not fully reflect non-European pedagogical traditions
  • Most task rubrics contain few negative criteria (which penalize undesirable behaviours), and so the influence of educationally inappropriate responses on the headline score is small.

For further detail on the tasks in the L2-Bench dataset, you can refer to our results paper linked in our previous results post.

Get involved

We’re releasing the L2-Bench dataset to the public because we believe strongly in the importance of stakeholder collaboration to create a common standard for assessing AI capabilities in language learning and teaching – we can’t do this alone.

Whether your interest is more technical or more pedagogical, we’d love for you to use it, test it, challenge it, and help us towards supporting an AI for Education evaluation ecosystem.

If you’d like to receive updates as the project develops, contribute to the next research phase, support our work, or leave feedback for us, please get in touch via the “Register Interest” form below.

AI in education is moving fast. Let’s make sure our evaluations keep up — and reflect the real work teachers do every day.