FI

Director, Research - AI Evals

Figma

Figma creates cloud-based design tools

San Francisco, CA • New York, NY • United Statesfull time
External · Company Career Page

This role is sourced from an external job board and is not a ProSculpt-verified employer. Review the listing carefully, never pay any fee to apply, and verify the company before sharing personal details.

Sign in to save Create account
Job Description

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Figma Research team is hiring an Director, Research - AI Evals to own how we measure the quality of Figma's AI-powered experiences. As Figma ships more AI capabilities across our products, the question "is this really good?" has never mattered more — and answering it rigorously is what this role exists to do. You'll define what "good" means for our AI features, build the frameworks and quality bars to measure it, and turn that into trusted signal that product teams rely on to decide what to ship. The ideal candidate brings deep, hands-on experience evaluating AI/LLM-powered products — blending human evaluation with automated, model-based approaches — along with the product instinct and communication skills to make evaluation genuinely useful. Partnering with Product, Design, Engineering, and Data Science, you'll sit upstream of nearly every AI shipping decision at Figma and directly shape the quality of features used by millions of people. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Own AI evaluation methods and operations for Figma's AI-powered experiences — define quality dimensions, design how we measure them, and turn results into decision-ready signal Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate Partner with engineering to stand up repeatable, reproducible evaluation pipelines and regression testing, so evaluation is a routine part of how AI features are built and shipped Produce clear readouts and dashboards that let stakeholders confidently make go/no-go and prioritization decisions Socialize a shared definition of quality so evaluation standards are adopted across teams rather than re-invented — and advocate for evaluation as a strategic partner in the product process Manage a small team to execute our AI evals in partnership with contractors, internal staff, and/or LLMs We'd love to hear from you if you have: 10+ years of experience in product, research, applied research, or a closely related field, including 2+ years of management experience Direct, hands-on experience owning the evaluation of AI/LLM-powered products Expertise designing and running AI evaluation — human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability — and sound judgment about when and how to apply automated/model-based approaches (e.g., LLM-as-judge), including their limitations Strength across both qualitative and quantitative methods, comfort with data and metrics, and the ability to reason about model behavior Demonstrated success in identifying the riskiest assumptions behind an ambiguous quality question, prioritizing them, and designing right-sized evaluation to build confidence A proven track record of gaining buy-in from executive and cross-disciplinary stakeholders — transcending methodology to articulate a larger user story and the "so what" to inspire action While it's not required, it's an added plus if you also have: Experience building or co-building automated evaluation pipelines and regression testing in partnership with engineering, or familiarity with eval tooling (e.g., Braintrust, LangSmith, DeepEval, or equivalents) Experience standing up a new function, practice, or discipline from scratch 2+ years in product design, user-centric product management, data science, product development, and/or front-end engineering A familiarity and depth of experience using Figma's products At Figma, one of our values is Grow as you go. We believe in hiring smart, curious people who are excited to learn and develop their skills. If you’re excited about this role but your past experience doesn’t align perfectly with the points outlined in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles. Pay Transparency Disclosure If based in Figma’s San Francisco or New York hub offices, this role has the annual base salary range stated below. Job level and actual compensation will be decided based on factors including, but not limited to, individual qualifications objectively assessed during the interview process (including skills and prior relevant experience, potential impact, and scope of role), market demands, and specific work location. The listed range is a guideline, and the range for this role may be modified. For roles that

Job Details
Posted
7/17/2026

Want a personalised feed?

Sign up to save jobs, track applications, and get AI-matched recommendations.

Create a free account