Table of Contents
A group of AI researchers has launched Sampura Research, a London-based nonprofit focused on improving how increasingly capable artificial intelligence systems are evaluated and supervised.
The organisation was launched by former Google DeepMind researchers Rishub Jain and Joshua Jacob, alongside co-founder Alex Adams. Sampura Research is focused on what it calls Human-AI Complementarity for Scalable Oversight, an approach that aims to combine the strengths of people and AI systems when evaluating advanced models.
Sampura Research and the Challenge of AI Oversight
The initial priority for Sampura Research is to build better AI “judges.”
The organisation defines a judge as a human, an AI system or a hybrid human-AI system that assesses the correctness and alignment of an AI system’s behaviour during a conversation or agent trajectory.
This research is intended to address a growing challenge in AI development. As models become more capable and autonomous, evaluating every output and action through human review alone can become difficult to scale. At the same time, relying entirely on AI systems to supervise other AI systems may introduce weaknesses that are difficult to identify.
Sampura Research’s approach is therefore centred on finding ways for humans and AI systems to contribute their complementary strengths to oversight.
Why Better AI Judges Matter
According to Sampura Research, imperfect evaluation systems can create problems throughout AI development.
For example, weak or incomplete evaluation methods can contribute to reward hacking, where an AI system learns to optimise for the evaluation process rather than genuinely achieving the intended objective. Imperfect judges can also lead to misleading research results and limit the effectiveness of AI monitoring during deployment.
The organisation believes that more reliable judges could improve several parts of the AI development process, including training data, reinforcement learning environments, red-teaming and model evaluations.
Building a Human-AI Evaluation Framework
Sampura Research plans to develop and maintain a broad leaderboard for evaluating AI judges.
The proposed system will include more than 20 sub-datasets covering areas such as:
- Deception detection
- Cultural bias
- Unsafe AI agent actions
- Alignment-related behaviour
The organisation also plans to test evaluation systems under situations where models are subject to optimisation pressure from the AI systems they are evaluating.
A key part of the research will involve determining how tasks should be divided between people and AI. Possible approaches include routing tasks to humans when an AI system has low confidence, training systems to decide whether a human or AI is better suited for a task, and breaking complex evaluations into smaller subtasks.
Sampura Research also plans to explore task-specific tools that can assist human evaluators.
Why Humans Could Remain Important
Sampura Research does not argue that humans should independently supervise every AI system. Instead, its research focuses on areas where humans may continue to provide value alongside increasingly capable AI models.
The organisation points to several possible advantages of Human-AI Complementarity, including differences between human and AI strengths, resilience against weaknesses in evaluation systems and the potential role of human judgement when new situations raise questions that existing AI systems may not have encountered.
This approach also differs from research strategies that primarily replace human evaluators with AI systems because AI-based evaluation can be faster and less expensive. Sampura Research wants to examine whether combining both forms of oversight can produce more reliable results.
Funding and Recruitment Plans
The nonprofit has received an initial $11 million grant from Coefficient Giving, including $7 million for its first year and a further $4 million pledged for future work.
Sampura Research is also recruiting founding technical researchers in London, with advertised salaries ranging from approximately £100,000 to £290,000. The organisation’s launch highlights the increasing competition for experienced researchers working in AI safety, evaluation and alignment.
What Comes After Building Better Judges?
After its initial work on evaluation systems, Sampura Research plans to study how improved judges could be deployed in AI training, evaluations and real-world monitoring.
Future research questions include whether stronger judges can reduce reward hacking, how human oversight can be balanced against speed and cost, and how evaluation systems should be updated as AI models evolve.
The organisation also intends to expand its work beyond classification-style evaluation. It plans to explore how instructions, specifications, rubrics and test cases can themselves be improved, since even an effective judge can produce poor results if it is evaluating against an incomplete or incorrect specification.
Conclusion
Sampura Research has entered the AI safety field with a specific focus on improving scalable oversight through Human-AI Complementarity.
Founded by former Google DeepMind researchers Rishub Jain and Joshua Jacob alongside Alex Adams, the London-based nonprofit plans to develop stronger AI judges, diverse evaluation datasets and methods for combining human and AI capabilities.
With $11 million in initial funding, Sampura Research is beginning with the challenge of building better evaluation systems before expanding its work toward AI deployment, monitoring and broader alignment research. Its central idea is not that humans should replace AI in oversight, or that AI should replace humans, but that identifying where their capabilities complement each other could become increasingly important as AI systems grow more powerful.

