We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

AI Safety Data Scientist

Skill
$76.00 - $82.00 / hr
401(k)
United States, New York, New York
Sep 13, 2026
Overview

Placement Type:

Temporary

Salary:

$76-82 Hourly

W2, Benefits and 401k matching

Start Date:

Sep 28, 2026

Location: Remote (US - East Coast / EST hours preferred)

Role Summary:

We are looking for an experienced Data Scientist to help us measure and monitor risks in conversational and agentic AI products. Reporting directly to the Data Science Lead, you will investigate emerging safety risks in production and turn them into scalable measurement, evaluation, and monitoring systems that inform policy, model, and product improvements.

What You'll Do:



  • Develop and scale risk monitoring and safety measurement systems across conversational, recommender, and tool-using AI features.
  • Design safety metrics and evaluation frameworks for production systems (including false positive/negative rates, safety risk prevalence, and decision rubrics).
  • Inspect, debug, and utilize Python data analysis scripts and SQL pipelines (leveraging LLM coding tools such as Claude or Codex to expedite analytical workflows).
  • Calibrate LLM-as-a-judge evaluation workflows, label evaluation datasets, and build reporting dashboards/visualizations for leadership and product stakeholders.
  • Work cross-functionally with Product, Engineering, and Trust & Safety teams to translate real-world production insights into safety mitigations and updated safety policies.
  • Communicate analytical findings and measurement metrics clearly to both technical and non-technical stakeholders.


Who You Are:



  • You have personally delivered safety evaluations, risk metrics, or mitigations for a real AI or machine-learning product.
  • You have strong Python and SQL data analysis skills, with the technical ability to independently inspect, critique, and debug pipeline code and queries.
  • You have hands-on experience designing evaluation datasets, rubrics, and measurement metrics.
  • You are comfortable structuring ambiguous problems and defining success criteria, methodologies, and tradeoffs with stakeholders.
  • You have experience working cross-functionally across multiple domains, including research, engineering, product, policy, or Trust & Safety.
  • You communicate clearly in writing and have a proven record of turning analytical research findings into clear product/policy action.
  • Minimum 2 years of direct AI safety data science or safety evaluation experience, OR 3 to 5+ years of broader Data Science experience if safety experience is foundational or project-based.


It's a Plus If You Have:



  • Experience calibrating LLM judges or building human-in-the-loop evaluations.
  • Experience evaluating multi-turn or tool-using agentic systems.
  • Experience with multilingual evaluation.

Applied = 0

(web-665cd84569-4dpxh)