Back to Home

About

The Science of Knowing
What AI Doesn't Know

Sonde AI builds benchmarks, training methods, and evaluation infrastructure that measure and improve AI's ability to recognize its own knowledge boundaries.

Why This Matters

The Problem
AI systems deployed in professional settings hallucinate when they lack information. In medicine, law, and engineering, confident wrong answers are worse than admitting uncertainty. Current models have no mechanism to pause and ask for clarification.
Our Approach
We treat question-asking as a measurable, trainable skill. Using controlled information masking and multi-turn Oracle interactions, we quantify how well models identify gaps and recover missing information through targeted questions.
Open Science
All our benchmarks, pipelines, and training code are open source. We believe advancing AI safety requires transparent, reproducible research that the entire community can build on.

What We Build

1

Evaluation Benchmarks

GDPval-AA covers 220+ professional tasks across 30 domains. HIL-Bench measures blocker detection quality with precision/recall metrics. Both use LLM judge panels for scalable, reproducible scoring.

2

Training Methods

We develop SFT and RL pipelines that teach models to ask better questions. Our published negative result (naive SFT degrades performance) and positive result (RL improves discrimination) guide the field toward effective approaches.

3

Infrastructure

E2B sandboxed execution, automated masking pipelines, multi-model comparison frameworks, and production-ready LoRA fine-tuning pipelines. Everything needed to evaluate and improve question-asking at scale.

Work With Us

We work with AI labs and enterprises to evaluate and improve model reliability in professional workflows.