User Researcher, AI Evaluations

Notion
Full-timeβ€’$196k-230k/year (USD)

πŸ“ Job Overview

Job Title: User Researcher, AI Evaluations

Company: Notion

Location: San Francisco, California, United States

Job Type: FULL_TIME

Category: User Experience Research / AI Product Evaluation

Date Posted: 2026-06-18

Experience Level: 5-10 Years

Remote Status: Hybrid (3 days in office)

πŸš€ Role Summary

  • Spearhead the definition and operationalization of evaluation frameworks and rubrics for Notion's cutting-edge AI-powered experiences, ensuring alignment with user expectations for helpfulness, trust, and transparency.

  • Conduct recurring and feature-specific user research studies, including qualitative interviews, surveys, and human-in-the-loop evaluations, to deeply understand user behavior and identify areas for improvement in AI interactions.

  • Translate complex qualitative insights into actionable, reusable measurement approaches and scoring guidelines that empower Product, Design, Engineering, and Data Science teams to consistently assess and enhance AI product quality.

  • Drive a systems-thinking approach to AI evaluation, anchoring research in real user workflows and jobs-to-be-done to ensure that evaluations reflect the end-to-end user journey, not just isolated feedback.

  • Collaborate closely with cross-functional partners to operationalize evaluation loops, integrating human judgment with automated or LLM-based assessment methods to create scalable and reliable quality assurance processes for AI features.

πŸ“ Enhancement Note: This role is a unique blend of deep UX research craft and the operationalization of evaluation methodologies, specifically tailored to the rapidly evolving landscape of AI products. The focus on "AI Evaluations" and "operationalizing insight into measurement" suggests a need for candidates who can bridge the gap between user understanding and scalable, data-driven assessment processes. This is not a purely generative AI research role, but one focused on evaluating the user experience of AI within a product context, requiring strong analytical, strategic, and collaborative skills.

πŸ“ˆ Primary Responsibilities

  • Develop and formalize comprehensive evaluation frameworks, rubrics, and scoring guidelines that define "good" AI experience quality, encompassing criteria such as helpfulness, trustworthiness, tone, control, and transparency.

  • Design and execute a variety of research methodologies, including longitudinal studies, feature-specific evaluations, side-by-side comparisons, and user interviews, to gather rich qualitative and quantitative data on AI product performance.

  • Synthesize research findings into actionable recommendations, clearly articulating failure modes, regressions, and opportunities for improvement to product and engineering teams.

  • Identify and analyze user breakdown and recovery behaviors within AI-powered workflows, translating these observations into practical guidance for system guardrails, UI improvements, and prioritization decisions.

  • Partner with Data Science and Engineering teams to explore and implement scalable evaluation loops, potentially including the calibration of LLM-as-judge approaches against human judgment and the development of "golden datasets."

  • Champion a user-centric approach to AI evaluation, ensuring that research efforts are grounded in real user scenarios, intents, and the complete interaction journey within Notion.

  • Facilitate cross-functional alignment on evaluation standards and foster a shared understanding of user expectations and AI product quality across Product, Design, Engineering, and Data Science.

  • Contribute to the selection and implementation of appropriate research and AI observability tooling to support efficient and effective evaluation processes.

πŸ“ Enhancement Note: The responsibilities highlight a strong emphasis on creating reusable assets (frameworks, rubrics) and embedding evaluation into ongoing product development cycles. This goes beyond typical ad-hoc research, requiring a strategic mindset focused on building scalable processes and measurement systems for AI product quality. The mention of "operationalize evaluation with partners" and "anchor evaluation in real workflows" points to a need for a candidate who can influence product strategy and collaborate effectively across technical disciplines.

πŸŽ“ Skills & Qualifications

Education:

  • Master’s or PhD in Human-Computer Interaction (HCI), Psychology, Behavioral Science, Anthropology, Sociology, or a closely related quantitative or qualitative field.

  • Alternatively: Equivalent practical experience demonstrating mastery in UX research and AI evaluation.

Experience:

  • 5+ years of dedicated UX research experience in an industry setting, with a proven track record of conducting impactful research leading to product improvements.

  • Demonstrated experience evaluating AI-enabled products, including LLMs, AI agents, or generative workflow automation tools.

  • Experience working collaboratively with Data Science and Machine Learning (ML) partners on measurement strategy and evaluation tooling. Required Skills:

  • Advanced UX Research Craft: Expertise in a range of qualitative (e.g., interviews, usability testing, ethnographic studies) and quantitative (e.g., surveys, benchmarking, experimental design) research methods, with the ability to select and apply the most appropriate techniques to complex problems.

  • AI Fluency & Systems Thinking: A strong understanding of AI concepts, particularly LLMs and AI agents, and the ability to reason about how model behavior, uncertainty, and system constraints impact user experience. Ability to think holistically about the end-to-end AI product experience.

  • Measurement & Rubric Development: Proven ability to translate abstract user expectations (e.g., trust, tone, usefulness, clarity) into concrete, measurable rubrics, scoring guidelines, and observable metrics.

  • Communication & Collaboration: Exceptional ability to communicate research insights clearly and persuasively to diverse audiences (Product, Design, Engineering, Data Science), build consensus, and drive actionable change. Experience managing stakeholders and aligning teams around shared definitions of quality.

  • Pragmatism & Ambiguity Tolerance: Demonstrated ability to thrive in fast-paced, ambiguous environments, prioritize effectively, and balance scrappy iteration with deep, rigorous investigations.

  • Workflow & Jobs-to-be-Done Focus: Experience grounding research in real user workflows, understanding user intent, and evaluating the full interaction journey from goal setting to iteration.

Preferred Skills:

  • Hands-on experience with LLM-as-judge methodologies, prompt design for evaluators, or the creation of "golden datasets" for AI evaluation.

  • Familiarity with AI research tooling for rapid synthesis (e.g., Dovetail, Listen Labs) and AI observability tooling (e.g., Braintrust).

  • Proficiency in data querying languages such as SQL, scripting languages like Python, or statistical software such as R, SAS, or Matlab.

  • Knowledge of the work and philosophies of computing pioneers like Douglas Engelbart, Alan Kay, or Bret Victor, and their relevance to human-computer interaction.

πŸ“ Enhancement Note: The emphasis on "operationalize insight into measurement" and "AI fluency and systems thinking" strongly suggests that candidates who can bridge the gap between qualitative user understanding and quantitative, scalable evaluation metrics will be highly valued. The preference for experience with specific AI evaluation methodologies and tooling indicates a sophisticated understanding of the current AI research landscape.

πŸ“Š Process & Systems Portfolio Requirements

Portfolio Essentials:

  • Demonstrate a portfolio showcasing projects where you have successfully translated user needs and qualitative insights into measurable outcomes or operationalized research findings into repeatable processes.

  • Include case studies that highlight your ability to define evaluation criteria, develop rubrics, and implement measurement strategies for complex product features, ideally involving AI or sophisticated technology.

  • Present examples of how you have collaborated with cross-functional teams (Product, Engineering, Data Science) to integrate research insights and evaluation processes into product development lifecycles.

  • Showcase projects where you have clearly articulated the impact of your research on product quality, user experience, or business objectives, with a focus on demonstrating ROI or efficiency gains. Process Documentation:

  • Provide evidence of your ability to design and document complex research workflows, from initial study design and participant recruitment to data synthesis, reporting, and follow-through on recommendations.

  • Illustrate your experience in creating and refining process documentation for research operations, including guidelines for study execution, data analysis, and team calibration.

  • Highlight instances where you have developed or improved processes for evaluating user experiences, particularly those involving AI, ensuring consistency, scalability, and actionable insights.

πŸ“ Enhancement Note: For a role focused on "evaluation operations," a portfolio is critical. It should not only showcase research craft but also the ability to build and scale evaluation processes. Candidates should prepare to present case studies that demonstrate their strategic thinking in defining what "good" looks like for AI products and how they operationalized that definition into repeatable evaluation methods.

πŸ’΅ Compensation & Benefits

Salary Range:

  • San Francisco/New York City: $196,000 - $230,000 USD per year.

  • Note: Compensation will be determined based on factors such as location, role scope, complexity, and candidate experience.

Benefits:

  • Competitive cash compensation.

  • Equity in the company.

  • Comprehensive benefits package (details not specified, but typical for tech companies). Working Hours:

  • 40 hours per week (standard full-time).

  • Hybrid work model requires 3 days per week in the San Francisco office (Anchor Days: Monday, Tuesday, Thursday), emphasizing in-person collaboration.

πŸ“ Enhancement Note: The salary range provided is specific to high-cost-of-living areas like San Francisco and New York City, reflecting a senior-level UX Researcher role with a specialized focus on AI. The hybrid work arrangement is a key differentiator, requiring candidates to be comfortable with in-office collaboration on designated days.

🎯 Team & Company Context

🏒 Company Culture

Industry: Software / Collaboration Tools / AI Workspace

Company Size: Notion is a rapidly growing technology company with a significant global presence, indicated by its substantial employee base and broad customer adoption. This size implies a dynamic environment with opportunities for impact and a focus on scalable solutions.

Founded: Founded in 2016, Notion is a relatively young company that has achieved significant market traction, suggesting a culture of rapid innovation, agility, and a strong vision for the future of work.

Team Structure:

  • The User Researcher will likely be part of a growing UX Research team, potentially embedded within product teams or operating as a central function supporting multiple product areas.

  • This role will report into a UX Research Lead or a Director of Product/Design, with a clear reporting line that offers guidance and mentorship.

  • Close collaboration is expected with Product Managers, Designers, Engineers, and Data Scientists, particularly those involved in developing and refining Notion's AI features. Methodology:

  • Notion emphasizes a "customer zero" approach, meaning employees actively use their own product, fostering a deep understanding and commitment to improving the user experience.

  • The company values "craft" and building "things that last," suggesting a commitment to quality, thoughtful design, and robust engineering.

  • There's a strong belief in the power of human collaboration, even within an AI-integrated workspace, as evidenced by the hybrid "Anchor Days" policy.

  • A data-driven approach is implied, especially with the focus on measurement and evaluation for AI products.

Company Website: https://www.notion.com/

πŸ“ Enhancement Note: Notion's culture appears to be a blend of high-growth tech company energy with a deliberate focus on user-centricity, product craft, and collaborative innovation. The "customer zero" philosophy and the emphasis on human collaboration within an AI-driven future are key cultural differentiators that candidates should understand and align with.

πŸ“ˆ Career & Growth Analysis

Operations Career Level: This is a Senior UX Researcher position, requiring 5+ years of experience. It's a specialized role focused on AI Evaluations, indicating a high level of expertise and the ability to operate with significant autonomy. The role is positioned to define and scale evaluation methodologies, suggesting a leadership trajectory within UX Research.

Reporting Structure: The role will likely report to a Director or Lead UX Researcher, providing mentorship and strategic direction. The researcher will also work very closely with cross-functional leads in Product Management, Design, and Engineering, influencing their roadmaps and development processes.

Operations Impact: The User Researcher will have a direct and substantial impact on the quality and user adoption of Notion's AI features. By defining what "good" looks like and operationalizing evaluation processes, this role will shape how millions of users perceive and trust AI within their workspace, ultimately influencing product strategy, user retention, and competitive positioning in the AI-powered productivity market.

Growth Opportunities:

  • Specialization: Deepen expertise in AI product evaluation, becoming a go-to expert within Notion and potentially the broader industry.

  • Leadership: Grow into a leadership role, managing a team of researchers focused on AI or product operations, or contributing to the strategic direction of the entire UX Research function.

  • Cross-functional Influence: Expand influence across Product, Engineering, and Data Science, driving strategic initiatives related to AI quality and user experience.

  • Methodological Innovation: Develop novel research methods and evaluation techniques tailored to the unique challenges of AI products.

πŸ“ Enhancement Note: This role offers significant growth potential for researchers interested in the intersection of UX, AI, and product operations. The opportunity to define and scale new evaluation processes for a core, strategic product area like AI is a strong indicator of future leadership and impact.

🌐 Work Environment

Office Type: Notion operates a hybrid work model, with designated "Anchor Days" (Monday, Tuesday, Thursday) for in-office collaboration. This indicates an office environment designed to foster teamwork, brainstorming, and spontaneous interaction.

Office Location(s): The role can be based in either San Francisco or New York City. These are major tech hubs, offering access to talent, industry events, and a vibrant professional community.

Workspace Context:

  • The hybrid model suggests a collaborative office space designed for team interaction, potentially featuring open areas, meeting rooms, and quiet zones.

  • As a tech company, expect access to modern tools, hardware, and software essential for research and product development.

  • Opportunities for informal "walk-bys" and close collaboration with Product Managers, Designers, and Engineers are inherent in the hybrid structure, fostering a dynamic and integrated work environment.

Work Schedule: The standard 40-hour work week is expected, with the flexibility to manage research projects effectively. The hybrid schedule requires presence in the office on Anchor Days, emphasizing the importance of in-person collaboration for specific team activities and knowledge sharing.

πŸ“ Enhancement Note: The hybrid work model is a significant aspect of the work environment. Candidates should be prepared for structured in-office collaboration days, which are integral to Notion's culture and operational strategy for AI development.

πŸ“„ Application & Portfolio Review Process

Interview Process:

  • Initial Screening: A recruiter will assess your resume and experience against the role requirements, focusing on your UX research background and any AI-related experience.

  • Hiring Manager Interview: A conversation with the hiring manager (likely a UX Research Lead or Director) to delve deeper into your experience, research philosophy, and fit with the team and Notion's culture.

  • Portfolio Review & Skill Assessment: A crucial stage where you'll present selected case studies from your portfolio. This will likely include a deep dive into your process, methodologies, and the impact of your work, especially any relevant to AI or complex system evaluations. Expect to discuss how you operationalize insights and define quality metrics.

  • Cross-functional Interviews: Meetings with key partners from Product, Design, and Engineering to assess your collaboration skills, ability to influence, and understanding of product development cycles.

  • Final Interview: Potentially with a more senior leader to discuss strategic alignment and overall fit.

Portfolio Review Tips:

  • Curate Strategically: Select 2-3 projects that best showcase your ability to define and operationalize evaluation frameworks, particularly for complex or AI-driven products. Highlight projects where you translated qualitative insights into quantitative measures or scalable processes.

  • Structure for Impact: For each project, clearly articulate the problem, your role, the methods used, the challenges faced (especially ambiguity or complexity), your specific contributions, the outcomes (quantifiable if possible), and the lessons learned.

  • Emphasize Operationalization: Clearly demonstrate how you moved beyond simple research findings to create reusable assets, processes, or frameworks that teams could adopt and use consistently. This is key for the "AI Evaluations" aspect.

  • AI Relevance: If you have direct AI evaluation experience, make it prominent. If not, draw parallels from complex system evaluations or projects requiring the definition of abstract quality criteria.

  • Tell a Story: Frame your presentations as narratives that highlight your problem-solving skills, strategic thinking, and ability to drive change through research.

Challenge Preparation:

  • Be prepared for a potential take-home assignment or a live exercise focused on evaluating a hypothetical AI feature, defining rubrics, or outlining an evaluation strategy.

  • Practice articulating your thought process clearly and concisely, as if explaining your approach to cross-functional partners.

  • Think about how you would balance rigor with speed in a fast-paced environment, and how you would define and measure success for an AI product.

πŸ“ Enhancement Note: The portfolio review is paramount. Candidates must demonstrate not just research expertise, but the ability to build and scale evaluation processes and metrics for AI products. This requires preparing case studies that highlight strategic thinking, operationalization, and cross-functional collaboration.

πŸ›  Tools & Technology Stack

Primary Tools:

  • UX Research Platforms: Experience with tools like Dovetail, Listen Labs, Maze, or Outset for qualitative data synthesis, user feedback collection, and analysis is highly desirable.

  • AI Observability Tools: Familiarity with platforms like Braintrust for monitoring and evaluating AI model performance and user experience.

  • Survey & Analytics Tools: Proficiency in standard tools for creating and distributing surveys, analyzing quantitative data, and potentially A/B testing.

Analytics & Reporting:

  • Data Querying: Experience with SQL is a significant plus for accessing and analyzing user data.

  • Scripting Languages: Familiarity with Python for data manipulation, analysis, or automation tasks.

  • Statistical Software: Knowledge of R, SAS, or Matlab can be beneficial for advanced quantitative analysis.

CRM & Automation:

  • While not primary for a UX Researcher, understanding how research feeds into CRM strategies or how automation impacts user workflows could be advantageous.

  • Familiarity with Notion's own platform as a user and potentially as a tool for organizing research findings.

πŸ“ Enhancement Note: The "Nice to Haves" section explicitly mentions specific AI research and observability tools. Candidates with experience in these or similar platforms will have a distinct advantage. Proficiency in SQL and Python is also highly valued for data-driven evaluation.

πŸ‘₯ Team Culture & Values

Operations Values:

  • Craft & Quality: A deep commitment to building high-quality, lasting products, which extends to the rigor and thoughtfulness applied to AI evaluations.

  • Human Collaboration: Valuing in-person interaction and collaborative problem-solving, especially during the designated "Anchor Days," to foster stronger teamwork and innovation.

  • Customer Zero: A culture of using their own product extensively to gain deep empathy and drive meaningful improvements, applying this mindset to understanding user interaction with AI.

  • Curiosity & Tinkering: An ethos of intellectual curiosity and a hands-on approach to exploring new technologies like AI, encouraging experimentation and discovery.

  • Impact Orientation: A focus on driving tangible product change and business outcomes through research and evaluation.

Collaboration Style:

  • Cross-functional Integration: The role is inherently collaborative, requiring seamless integration with Product, Design, Engineering, and Data Science teams. Expect frequent interaction and joint problem-solving.

  • Data-Informed Decision Making: A culture that values data and research insights as critical inputs for product strategy and development.

  • Iterative & Agile: Embracing an iterative approach to product development and research, balancing quick feedback loops with deeper investigations.

  • Open Communication: An environment that encourages open dialogue, feedback exchange, and knowledge sharing to foster continuous improvement.

πŸ“ Enhancement Note: Notion's emphasis on "craft," "human collaboration," and "customer zero" suggests a culture that is both demanding of high-quality work and supportive of teamwork and deep product understanding. Candidates should be prepared to thrive in an environment that values both individual expertise and collective effort.

⚑ Challenges & Growth Opportunities

Challenges:

  • Defining "Good" AI: Establishing objective, universally applicable criteria for evaluating AI experiences, which are often subjective and context-dependent.

  • Scalability of Evaluation: Developing evaluation methods that can keep pace with the rapid iteration cycles common in AI product development and scale across multiple product areas.

  • Balancing Rigor and Speed: Navigating the tension between conducting thorough, in-depth research and delivering timely insights in a fast-moving AI product environment.

  • Cross-functional Alignment: Ensuring consistent understanding and application of evaluation rubrics and insights across diverse teams with varying priorities and technical backgrounds.

  • Measuring Trust & Subjectivity: Quantifying subjective user experiences like trust, helpfulness, and tone in AI interactions.

Learning & Development Opportunities:

  • AI Research Specialization: Becoming a leading expert in the emerging field of AI product evaluation, including LLM assessment methodologies.

  • Methodological Innovation: Opportunity to develop and pioneer new research techniques tailored for AI-driven products.

  • Strategic Influence: Gaining significant influence over the direction and quality of Notion's AI product roadmap.

  • Cross-functional Skill Development: Enhancing collaboration and communication skills by working closely with technical and product leadership.

  • Industry Exposure: Engaging with cutting-edge AI research and tooling, potentially attending relevant conferences or workshops.

πŸ“ Enhancement Note: The primary challenge lies in operationalizing the evaluation of AI, a complex and rapidly evolving domain. Success in this role will require strong adaptability, strategic thinking, and the ability to build scalable processes in an ambiguous environment.

πŸ’‘ Interview Preparation

Strategy Questions:

  • "How would you define and measure 'trust' in an AI-powered workspace tool like Notion?"

  • "Describe a framework you would use to evaluate the helpfulness and accuracy of an AI agent's response within a user's workflow."

  • "How would you operationalize AI evaluation rubrics so that Product Managers and Engineers can use them consistently without direct researcher involvement?"

  • "What are the key differences in evaluating generative AI features versus traditional software features, and how would you adapt your research approach?"

  • "How would you identify and mitigate potential biases in AI outputs that could negatively impact user experience?" Company & Culture Questions:

  • "What excites you about Notion's mission and its approach to AI?"

  • "How do you see yourself contributing to Notion's 'customer zero' culture?"

  • "Describe your ideal collaboration dynamic with Product Managers, Designers, and Engineers, especially regarding AI features."

  • "How do you handle ambiguity and fast-paced development cycles in your research work?"

  • "What are your thoughts on hybrid work and in-office collaboration days?" Portfolio Presentation Strategy:

  • Focus on Process & Impact: For each case study, clearly articulate the "why" (problem), "how" (your process and methodology), and "what" (outcomes and impact).

  • Highlight Operationalization: Explicitly demonstrate how your research led to the creation of scalable processes, frameworks, or metrics that teams could adopt. Use phrases like "developed a reusable rubric," "implemented a longitudinal tracking system," or "operationalized feedback loops."

  • Quantify When Possible: If you can, present metrics that show the impact of your work (e.g., improvement in user satisfaction scores, reduction in reported issues, increased adoption of AI features).

  • Address AI Nuances: If presenting non-AI projects, draw parallels to how your approach to defining abstract quality or managing complex systems could apply to AI evaluation.

  • Be Ready for Deep Dives: Anticipate detailed questions about your methodology, decision-making process, and how you handled challenges or disagreements.

πŸ“ Enhancement Note: Interview preparation should heavily emphasize the "operationalization" aspect of the role. Candidates need to demonstrate not just research prowess but the ability to build systems and metrics for evaluating AI quality, and how they would integrate these into Notion's product development lifecycle.

πŸ“Œ Application Steps

To apply for this operations position:

  • Submit your application through the provided link on Ashby.

  • Tailor your Resume: Highlight your 5+ years of UX research experience, specifically emphasizing any work with AI products, complex systems, or the development/operationalization of evaluation frameworks and metrics. Use keywords from the job description such as "AI evaluation," "rubrics," "measurement strategy," "systems thinking," and "cross-functional collaboration."

  • Prepare Your Portfolio: Curate 2-3 key projects that best demonstrate your ability to define and scale evaluation processes for user experiences, particularly those involving AI or abstract quality criteria. Be ready to present your process, outcomes, and the impact of your work, focusing on how you translated insights into actionable, repeatable measurement.

  • Research Notion: Understand Notion's product, mission, and its growing role in the AI-powered workspace. Familiarize yourself with their public statements on AI and collaboration.

  • Practice Interview Responses: Prepare for behavioral questions by using the STAR method (Situation, Task, Action, Result) and focus on examples that showcase your strategic thinking, collaboration skills, and experience with ambiguity and fast-paced environments. Practice articulating your approach to AI evaluation challenges.

⚠️ Important Notice: This enhanced job description includes AI-generated insights and operations industry-standard assumptions. All details should be verified directly with the hiring organization before making application decisions.


Application Requirements

Requires 5+ years of industry UX research experience with a strong ability to translate user expectations into concrete metrics. Candidates should possess AI fluency and experience evaluating LLM-enabled products or generative workflows.