AI Benchmark Quality Reviewer
Braintrust
Job description
About the role
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.
Key responsibilities
- Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
- Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
- Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
- Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.
- Write concise, evidence‑backed findings and verify that revisions address the issues found.
Required profile
- At least five years of relevant technical or analytical experience.
- Ability to read Python, SQL, shell scripts, structured data and execution logs.
- Strong written English, analytical judgment and attention to detail.
- Ability to give specific, reproducible feedback and explain uncertainty clearly.
- Experience in AI evaluation, technical QA, data analysis or benchmark development is helpful but not required.
Required skills
- Python
- SQL
- Shell scripting
- Structured data analysis
- Execution‑log interpretation
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Paraguay.
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 6 days ago
Expires 1 month from now
31 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Braintrust