Case Studies
Selected projects in AI data operations and evaluation.
Improving Human Judgment for AI Safety Evaluation
Challenge
Ambiguous policy guidelines led to inconsistent human judgments.
Approach
--> Human review
--> Ground truth
--> Guideline refinement
Outcome
--> More consistent annotations
--> Better evaluation data
Key Insight
Reliable evaluation depends on consistent human judgment-not just good policies.
Designing a Searchable Job Title Taxonomy
Challenge
Job titles meant different things across different companies.
Approach
--> Taxonomy
--> Ontology
--> Search design
Outcome
--> Better discovery
--> Cleaner analytics
Key Insight
The best taxonomies make real user tasks easier.
Combining LLMs with Deterministic Validation
Challenge
LLM automation wasn't reliable enough by itself.
Approach
--> Gemini
--> Validation
--> Quality checks
Outcome
--> More structured data
--> Less manual work
Key Insight
Automation works best when paired with deterministic validation.