AI Benchmark Quality Reviewer
Braintrust
Descripcion del puesto
About the role
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.
Key responsibilities
- Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
- Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
- Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
- Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.
- Write concise, evidence‑backed findings and verify that revisions address the issues found.
Required profile
- At least five years of relevant technical or analytical experience.
- Ability to read Python, SQL, shell scripts, structured data and execution logs.
- Strong written English, analytical judgment and attention to detail.
- Ability to give specific, reproducible feedback and explain uncertainty clearly.
- Experience in AI evaluation, technical QA, data analysis or benchmark development is helpful but not required.
Required skills
- Python
- SQL
- Shell scripting
- Structured data analysis
- Execution‑log interpretation
Questions fréquentes
Por que reporta esta oferta?
Explorar más
Salarios, guías y búsquedas en Paraguay.
Postula en 30 segundos
Ingresa tu email para postular. Se creara una cuenta automaticamente.
Al continuar, aceptas nuestras condiciones de uso.
Ya tienes cuenta? Iniciar sesion
Publicado hace 2 días
Expira en 1 mes
24 vistas · 0 interested
Aumenta tus posibilidades
Sube tu CV: te propondremos las ofertas que coinciden con tu perfil.
Analizando tu CV...
Braintrust