InterviewDB
Experience
·
Paris
LLM Output Annotator: Build a Tool That Labels LLM Responses with Quality Tags
phone
Interview Experience
Problem You are building an annotation tool that takes an LLM-generated response and a reference answer, and labels the response with quality tags from a predefined set: ["correct", "partially_correct", "incorrect", "hallucinated", "refused", "off_topic"]. Implement the annotation logic using simple heuristics (exact match, keyword overlap, refusal detection). Example: Follow-ups What NLP metric (BLEU, ROUGE, BERTScore) gives a better signal than keyword overlap for partial correctness? How woul…
Full Details
🔒
Unlock all Harvey questions
Full insider details, leaked discussions, and candidate experiences.
Get full access — $100 a year, unlimited accessAbout This Question
This is a candidate experience report from a harvey interview during the phone round.