Product Designer, Evals & Prompts

Anthropic
Full-timeβ€’$305k-385k/year (USD)β€’San Francisco, United States

πŸ“ Job Overview

Job Title: Product Designer, Evals & Prompts

Company: Anthropic

Location: San Francisco, CA

Job Type: Full-Time

Category: Product Design / AI Operations

Date Posted: 2026-09-04

Experience Level: Mid-Senior (5-10 years)

Remote Status: Hybrid

πŸš€ Role Summary

  • Designs and refines prompts for AI features and behaviors, ensuring alignment with product strategy and user expectations.

  • Develops and maintains robust evaluation pipelines for Large Language Models (LLMs) to assess prompt effectiveness and model performance.

  • Builds user-friendly, low-code internal tools that empower designers to conduct evaluations without engineering support.

  • Supports critical model release cycles by rigorously testing product surfaces and implementing necessary prompt fixes and migrations.

  • Contributes to the scaling and reliability of the eval harness, ensuring consistent and trustworthy testing environments across various models and tools.

πŸ“ Enhancement Note: This role is crucial for ensuring the reliability, safety, and user-centricity of Anthropic's AI products, specifically Claude. It bridges the gap between product design, prompt engineering, and rigorous evaluation methodologies, requiring a blend of technical proficiency and a deep understanding of user experience in AI interactions. The emphasis on building tools for designers highlights a strategic focus on democratizing evaluation processes.

πŸ“ˆ Primary Responsibilities

  • Write, test, and iterate on prompts for Claude's tools, features, and core behaviors across various product surfaces, ensuring intended user experiences are delivered.

  • Develop automated evaluation graders and rubrics, transforming manual assessment processes into scalable, reproducible testing frameworks.

  • Design and implement visual, low-code eval tools that enable product designers to easily assemble comparison sets, grade prompt variants, and analyze results independently.

  • Conduct thorough transcript analysis to identify nuances and limitations missed by automated evaluations, feeding insights back into prompt and eval design.

  • Support model release processes by performing regression testing on all product surfaces, writing prompt fixes, and developing prompts for new features with data-driven insights.

  • Establish and maintain the integrity of the eval harness, ensuring a stable test environment for 50-100 tools and accurately diagnosing regressions between the harness and the model.

  • Package actionable feedback from evaluations (e.g., graders, human feedback questions, preference pairs) for model training when prompt-based solutions are insufficient.

πŸ“ Enhancement Note: The responsibilities indicate a hands-on role requiring deep engagement with both the creative aspects of prompt design and the rigorous, analytical side of LLM evaluation. The emphasis on "Production-quality Python" and "building internal tools with a real interface" suggests a strong software engineering component within the product design function.

πŸŽ“ Skills & Qualifications

Education: Bachelor’s degree or an equivalent combination of education, training, and/or experience in a relevant field.

Experience: Years of experience will correlate with internal job level requirements; typically 5-10 years of experience is expected for this level of responsibility.

Required Skills:

  • Production-quality Python programming expertise for building robust applications and pipelines.

  • Proven experience in designing, building, and maintaining evaluation pipelines for LLM products, including graders, rubrics, comparison sets, and regression suites.

  • Demonstrated ability to create internal tools with intuitive user interfaces for non-technical users (designers).

  • Experience in setting up test harnesses, sandboxing tool calls, and managing comparable test run configurations.

  • Practical understanding of shipping prompts, including awareness of model-specific prompt performance variations and failure modes.

  • Proficiency in analyzing raw model transcripts to extract qualitative insights beyond quantitative scores. Preferred Skills:

  • Direct experience working within a model-launch cycle and understanding the associated evaluation demands.

  • Familiarity with A/B testing methodologies and the ability to correlate offline evaluation results with online product performance.

  • Experience with front-end development or notebook-to-app transitions, with strong opinions on effective data visualization for evaluation results.

  • Proven ability to translate product rubrics into effective training signals, such as graders, human feedback questions, or preference pairs.

  • A user-centric mindset, prioritizing the actual behavior and experience of Claude for end-users over solely focusing on metric movement.

πŸ“ Enhancement Note: The qualifications highlight a unique blend of software engineering, AI/ML evaluation, and product design sensibilities. The emphasis on "reads transcripts, not only scores" is a critical differentiator, indicating a need for qualitative analytical skills alongside quantitative rigor.

πŸ“Š Process & Systems Portfolio Requirements

Portfolio Essentials:

  • Showcase examples of designed prompts and their impact on AI behavior, ideally with quantifiable results or qualitative user feedback.

  • Include case studies demonstrating the development and implementation of LLM evaluation pipelines, highlighting methodologies, tools used, and outcomes.

  • Present examples of internal tools or dashboards created for non-technical users to facilitate data analysis or process management.

  • Demonstrate experience in setting up test harnesses or sandboxing environments, illustrating your approach to ensuring reliable and reproducible testing.

  • Provide evidence of contributions to model release cycles, detailing how you validated product surfaces and addressed regressions or implemented prompt fixes. Process Documentation:

  • Document the design and iteration process for complex prompts, including rationale, testing strategies, and refinement based on evaluation feedback.

  • Outline the methodology for building automated evaluation systems, from rubric definition to grader implementation and pipeline orchestration.

  • Detail the process of developing and deploying internal tools, emphasizing user feedback incorporation and iterative improvement for usability.

  • Describe your approach to analyzing model output and transcripts, including how you derive actionable insights for prompt engineering and model training.

πŸ“ Enhancement Note: For this role, a portfolio should heavily emphasize practical application of prompt engineering and evaluation skills. Candidates should be prepared to walk through the technical details of their evaluation harnesses, the logic behind their prompt designs, and the user-centricity of their internal tools.

πŸ’΅ Compensation & Benefits

Salary Range: $305,000 - $385,000 USD per year.

Benefits:

  • Competitive compensation package.

  • Optional equity donation matching.

  • Generous vacation policy.

  • Comprehensive parental leave.

  • Flexible working hours.

  • Potential visa sponsorship for qualified candidates.

  • Hybrid work environment with at least 25% in-office time.

Working Hours: Standard 40 hours per week, with flexibility.

πŸ“ Enhancement Note: The provided salary range is at the higher end, reflecting the specialized nature of AI product design and evaluation, the demand for experienced professionals in this field, and the cost of living in San Francisco. The benefits package is comprehensive, aligning with industry standards for leading tech companies. The hybrid policy with a minimum of 25% in-office time suggests a need for in-person collaboration, particularly for roles involving complex product development and team alignment.

🎯 Team & Company Context

🏒 Company Culture

Industry: Artificial Intelligence (AI) Research and Development, AI Safety.

Company Size: Approximately 251-500 employees (based on general knowledge of Anthropic's growth phase).

Founded: 2021. Anthropic was founded by former members of OpenAI, with a strong focus on AI safety and creating beneficial AI systems. This history informs a culture deeply rooted in research, rigorous scientific inquiry, and a commitment to ethical AI development.

Team Structure:

  • The role sits within the Product Prompt and Eval Design team, which is part of the broader Product Design organization.

  • Day-to-day collaboration occurs with surface owners (Product Managers/Designers) and engineers within various product teams.

  • Close partnership with the prompt engineering team is expected, especially during model releases.

  • The eval harness team is a key functional area this role will contribute to and potentially lead. Methodology:

  • Data-driven decision-making, heavily relying on empirical evidence from evaluations and A/B tests to guide prompt design and product strategy.

  • Emphasis on scientific rigor, treating AI development as an empirical science akin to physics or biology.

  • Collaborative research and development approach, with frequent discussions to align on high-impact work.

  • Iterative development cycles for prompts and evaluations, incorporating feedback from both automated systems and human analysis.

Company Website: https://www.anthropic.com/

πŸ“ Enhancement Note: Anthropic's mission of creating safe and beneficial AI systems suggests a culture that values deep thinking, ethical considerations, and meticulous execution. The "big science" approach implies a focus on long-term impact and tackling complex, large-scale problems rather than incremental improvements.

πŸ“ˆ Career & Growth Analysis

Operations Career Level: This role is positioned as a foundational member of the eval side of the Product Prompt and Eval Design team, indicating a mid-to-senior level position. It requires significant autonomy and the ability to build critical infrastructure (eval harness, low-code tools). The scope involves direct impact on product quality, safety, and user experience.

Reporting Structure: The role reports to a lead within the Product Prompt and Eval Design team, likely a Product Design Lead or a similar senior role overseeing AI product evaluation and prompt strategy. Collaboration is cross-functional, involving close work with product engineers, prompt engineers, and other designers.

Operations Impact: This role has a direct and significant impact on the reliability, safety, and perceived quality of Anthropic's AI products, particularly Claude. By ensuring prompts are effective and behaviors are rigorously evaluated, the role directly influences user trust, product adoption, and adherence to AI safety principles. The development of evaluation tools also scales the impact across multiple product surfaces and teams.

Growth Opportunities:

  • Specialization: Deepen expertise in LLM evaluation methodologies, prompt engineering, AI safety, and human-computer interaction for AI.

  • Leadership: Grow into a lead role within the Product Prompt and Eval Design team, managing evaluation infrastructure, mentoring junior team members, and shaping strategic direction for eval methodologies.

  • Cross-functional Mobility: Transition into broader product management, AI research, or specialized prompt engineering roles within Anthropic, leveraging a deep understanding of model behavior and user interaction.

  • Tool Development Expertise: Become a go-to expert for building internal developer and designer tools within an AI research environment.

πŸ“ Enhancement Note: The role offers a unique opportunity to shape the future of AI product evaluation and design at a leading AI safety company. The combination of technical depth and product focus provides a strong foundation for a career in applied AI.

🌐 Work Environment

Office Type: Hybrid work model, with a requirement to be in the office at least 25% of the time. Anthropic's office is located in San Francisco, CA, suggesting a modern, collaborative workspace designed for focused work and team interaction.

Office Location(s): San Francisco, CA. Specific details about office amenities and accessibility would be available upon inquiry or during the interview process.

Workspace Context:

  • Collaborative environment fostering open communication and knowledge sharing among researchers, engineers, and designers.

  • Access to cutting-edge AI research and development resources, including advanced computing infrastructure.

  • A culture that supports empirical science and data-driven insights, encouraging experimentation and learning.

  • Opportunities to work closely with world-class AI talent on some of the most pressing challenges in AI safety and development.

Work Schedule: Flexible working hours are offered, supporting a healthy work-life balance while ensuring the necessary collaboration for a hybrid and fast-paced environment.

πŸ“ Enhancement Note: The hybrid model implies a need for individuals who can manage their time effectively, contributing both independently and collaboratively. The emphasis on collaboration suggests opportunities for direct interaction with diverse teams, which is crucial for understanding and refining AI product behaviors.

πŸ“„ Application & Portfolio Review Process

Interview Process:

  • Initial Screen: A recruiter or hiring manager will review your application and conduct an initial screening call to assess your background and fit.

  • Technical/Design Interviews: Expect multiple rounds focusing on your Python proficiency, LLM evaluation experience, prompt design principles, and internal tool development capabilities. This may include coding exercises and discussions on past projects.

  • Portfolio Review: A dedicated session to present and discuss your portfolio. Be prepared to deep-dive into specific projects, explaining your process, technical choices, and the impact of your work.

  • Team/Cross-functional Interviews: Discussions with potential teammates and stakeholders (e.g., prompt engineers, product managers, other designers) to assess collaboration style, problem-solving approach, and cultural fit.

  • Hiring Manager Interview: A final discussion to cover overall fit, career aspirations, and alignment with Anthropic's mission and values.

Portfolio Review Tips:

  • Showcase Impact: For each project, clearly articulate the problem, your solution (prompt design, eval pipeline, tool), and the measurable impact (e.g., improved model performance, reduced errors, increased designer efficiency, successful model release).

  • Technical Depth: Be ready to discuss the architecture of your evaluation harnesses, the logic of your graders, and the specific Python libraries or frameworks you used. Explain why you made certain technical decisions.

  • User-Centricity: For any tools you've built, demonstrate how they simplify complex processes for designers. Highlight user feedback and how it influenced your design.

  • Prompt Examples: Include examples of prompts you've written or refined, explaining the rationale behind their structure, parameters, and how they were tested and iterated upon.

  • LLM Evaluation Focus: Prioritize projects that showcase your ability to build and maintain evaluation pipelines, including experience with regression testing and transcript analysis.

Challenge Preparation:

  • Coding Challenges: Practice Python coding problems, focusing on data structures, algorithms, and practical application relevant to data processing and pipeline building.

  • System Design: Be prepared to discuss how you would design an evaluation harness or a low-code tool for a given scenario. Think about scalability, maintainability, and user experience.

  • Problem-Solving Scenarios: Anticipate questions about how you would debug a failing prompt, design an evaluation for a new AI feature, or improve an existing evaluation process.

πŸ“ Enhancement Note: The interview process is designed to thoroughly assess both technical aptitude and practical application in the specific domain of LLM evaluation and prompt design. A strong portfolio demonstrating hands-on experience and clear articulation of impact will be critical.

πŸ›  Tools & Technology Stack

Primary Tools:

  • Python: The core language for production-quality code, pipeline development, and internal tool creation. Proficiency is non-negotiable.

  • LLM Frameworks/Libraries: Experience with libraries for interacting with LLMs (e.g., OpenAI API, Hugging Face Transformers, Anthropic's own SDKs) is expected.

  • Data Analysis & Visualization: Tools like Pandas, NumPy, Matplotlib, Seaborn for data manipulation and plotting within evaluation pipelines.

  • Web Frameworks (for internal tools): Familiarity with frameworks like Flask, Django, or Streamlit for building user interfaces for designers.

Analytics & Reporting:

  • Experimentation Platforms: Experience with A/B testing frameworks to connect offline evals to online outcomes.

  • Dashboarding Tools: Potentially tools like Tableau, Looker, or custom-built dashboards for visualizing evaluation results and performance metrics.

  • Logging & Monitoring: Tools for tracking eval harness performance and identifying regressions.

CRM & Automation:

  • Version Control: Git is standard for code management.

  • CI/CD: Familiarity with continuous integration and continuous deployment pipelines for tool and harness development.

  • Cloud Platforms: Experience with cloud services (AWS, GCP, Azure) for hosting infrastructure and running pipelines.

πŸ“ Enhancement Note: The emphasis on "production-quality Python" and "building internal tools" suggests a need for strong software engineering skills applied within an ML/AI context. The role requires the ability to not only use existing tools but also to build new ones to solve specific evaluation and design challenges.

πŸ‘₯ Team Culture & Values

Operations Values:

  • Safety & Beneficence: A deep commitment to building AI systems that are safe, reliable, and beneficial for humanity, underpinning all design and evaluation efforts.

  • Empirical Rigor: A strong belief in data-driven decision-making and scientific methodology, using evidence from evaluations to guide product development.

  • Collaboration & Openness: A culture of intense collaboration and open communication, valuing diverse perspectives and collective problem-solving.

  • Impact-Oriented: A focus on achieving significant, long-term impact rather than working on smaller, isolated problems.

  • User-Centricity: A dedication to understanding and meeting user needs and expectations, ensuring AI interactions are intuitive and helpful.

Collaboration Style:

  • Highly collaborative, working closely with product managers, engineers, and other designers.

  • Emphasis on clear communication, especially regarding technical findings, prompt behavior, and evaluation results.

  • Iterative feedback loops, with regular reviews of prompts, evaluations, and tools to ensure continuous improvement.

  • A willingness to share knowledge and mentor others, particularly in democratizing evaluation processes.

πŸ“ Enhancement Note: Anthropic's culture is deeply intertwined with its mission. Candidates should demonstrate a genuine passion for AI safety and a collaborative spirit, reflecting the company's "big science" approach.

⚑ Challenges & Growth Opportunities

Challenges:

  • Scaling Evaluation: Developing evaluation systems that can keep pace with rapid model development and a growing number of product features and tools.

  • Bridging Offline/Online: Effectively translating the insights from offline evaluations into measurable improvements in online A/B tests and user experiences.

  • Democratizing Tools: Creating internal tools that are truly user-friendly for designers who may not have deep technical backgrounds, while maintaining robustness.

  • Navigating Model Changes: Adapting prompts and evaluations to frequent model updates, understanding how changes in model capabilities affect existing behaviors.

  • Balancing Safety and Capability: Ensuring that prompt designs and evaluations uphold safety standards without overly restricting the model's helpfulness or capabilities.

Learning & Development Opportunities:

  • Cutting-Edge AI Research: Direct exposure to and involvement in state-of-the-art AI research and development.

  • Specialized Skill Development: Opportunities to deepen expertise in LLM evaluation, prompt engineering, AI safety mechanisms, and human-AI interaction design.

  • Tool Development: Gaining experience in building sophisticated internal developer and designer tools within a fast-paced AI research environment.

  • Industry Leadership: Potential to contribute to shaping best practices in LLM evaluation and prompt design within the broader AI community.

πŸ“ Enhancement Note: The challenges presented are inherent to working at the forefront of AI development. Success in this role requires adaptability, a strong problem-solving mindset, and a commitment to continuous learning and innovation.

πŸ’‘ Interview Preparation

Strategy Questions:

  • "Describe your process for evaluating the performance of an LLM's responses for a specific product feature. How would you translate your findings into actionable prompt improvements?" (Focus on your evaluation pipeline, rubric design, and iterative prompt refinement).

  • "Imagine you need to build a tool for designers to compare prompt variants across different models. What would be your key design considerations for usability and effectiveness?" (Highlight your experience with low-code tools and user empathy).

  • "How would you approach setting up a test harness to ensure reliable regression testing for 50+ AI tools across multiple model versions?" (Discuss your experience with sandboxing, configuration management, and ensuring test stability). Company & Culture Questions:

  • "Why are you interested in Anthropic's mission of AI safety and beneficence? How does that align with your career goals?" (Show genuine interest and understanding of the company's core values).

  • "Describe a time you had to collaborate closely with engineers or product managers to ship a complex feature or fix a critical issue. What was your role and how did you ensure alignment?" (Emphasize communication, teamwork, and problem-solving).

  • "How do you balance the need for AI to be helpful and capable with the imperative for it to be safe and aligned?" (Discuss your understanding of AI safety principles and trade-offs). Portfolio Presentation Strategy:

  • Structure Your Narrative: For each project, clearly define the problem, your specific contribution, the solution/methodology, and the quantifiable or qualitative outcome.

  • Highlight Technical Aspects: Be prepared to walk through code snippets (if applicable), discuss your architecture choices for pipelines and tools, and explain the logic behind your evaluation metrics.

  • Focus on Impact: Quantify your achievements whenever possible. If direct metrics are unavailable, focus on how your work improved efficiency, reduced errors, or enhanced user experience.

  • Demonstrate User Empathy: For tool-building projects, explain how you considered the end-user (designers) and how their feedback shaped the final product.

  • Be Ready for Deep Dives: Anticipate questions about the nuances of your projects, alternative approaches you considered, and challenges you overcame.

πŸ“ Enhancement Note: Interview preparation should focus on demonstrating a strong understanding of LLM evaluation, practical software engineering skills in Python, and a user-centric approach to product design within the AI domain. Your portfolio is your primary tool for showcasing these capabilities.

πŸ“Œ Application Steps

To apply for this operations position:

  • Submit your application through the provided link on Greenhouse.

  • Tailor Your Resume: Highlight specific experience with Python, LLM evaluation pipelines, prompt engineering, internal tool development, and A/B testing. Use keywords from the job description.

  • Curate Your Portfolio: Select 2-3 key projects that best showcase your ability to design prompts, build evaluation systems, and create user-friendly tools. Prepare to discuss them in detail.

  • Prepare Your Narrative: Practice articulating your experience and impact concisely, focusing on the "what," "how," and "why" of your work, especially concerning AI product quality and safety.

  • Research Anthropic: Understand their mission, recent research, and the role of Claude. Be ready to discuss how your skills contribute to their goals.

⚠️ Important Notice: This enhanced job description includes AI-generated insights and operations industry-standard assumptions. All details should be verified directly with the hiring organization before making application decisions.


Application Requirements

Candidates must have production-quality Python skills and extensive experience building evaluation pipelines for LLM products. A bachelor's degree or equivalent experience is required, along with the ability to build internal tools and analyze model transcripts.