← All opportunities

Senior Data Scientist for Copilot Evals

Microsoft · United States, Washington, Redmond

Build and operate a data-driven prioritization framework for CADET quality forums and partner teams that combines DSAT, customer impact, usage, severity, strategic importance, representation, and current evaluation coverage to guide quality investments. Define and maintain intent, sub-intent, customer scenario, coverage, and loss-pattern taxonomies that can be used consistently across signal intake, triage, evaluation, and reporting. Measure how well evaluation portfolios represent production traffic, customer segments, workflow complexity, locales, grounding paths, and material failure modes. Identify underrepresented customers, intents, scenarios, and loss patterns, and translate those gaps into evaluation and data-collection priorities. Design quality gates for top intents and top customers, including success thresholds, segmentation, run cadence, escalation criteria, and reporting. Build recurring scorecards that connect offline evaluation movement with online measures such as DSAT, task completion, retries, abandonment, and escalation; detect meaningful quality changes; and alert accountable owners when action is required. Analyze offline-online agreement, evaluation freshness, regression coverage, grader reliability, and quality movement over time. Develop sampling, weighting, deduplication, clustering, and trend-detection approaches for noisy customer and product signals using resource- and performance-optimized data-analysis solutions that make effective use of CPU, GPU, and platform capacity. Use causal and experimental methods where appropriate to distinguish correlation, attribution, and treatment impact, and translate the findings into concrete product, model, data, and evaluation investment decisions. Partner to translate customer evidence into valid task distributions, datasets, metrics, and reward signals and to encode metrics, taxonomies, data-quality checks, and reporting into automated pipelines. Leverage team signals from customer engagements to understand workflows, business impact, expected outcomes, and gaps hidden by aggregate metrics; produce clear recommendations for product, model, data, and evaluation investments; and communicate them to senior leaders. Establish solid practices for data provenance, privacy, responsible use, reproducibility, and metric governance. Doctorate in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 1+ year(s) data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results) OR Master's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 3+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results) OR Bachelor's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 5+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results) OR equivalent experience. Solid experience using data to shape product strategy and decisions in a complex, high-scale product or platform environment. Expertise in SQL and at least one analytical programming language such as Python or R. Solid foundation in statistical analysis, experimentation, sampling, segmentation, measurement, and data visualization. Experience integrating noisy quantitative and qualitative signals into actionable prioritization or measurement frameworks. Ability to define durable metrics and taxonomies, explain their limitations, and prevent misleading interpretation. Demonstrated ability to communicate complex analysis clearly to technical, product, and executive audiences. Experience with AI product quality, LLM or agent evaluation, DSAT or customer feedback analysis, experimentation, or model telemetry. Experience designing representative datasets, coverage models, quality scorecards, or regression portfolios. Experience with clustering, text analytics, embeddings, classification, anomaly detection, or other methods for mining unstructured feedback. Experience connecting offline evaluation results with online product and customer outcomes. Experience working directly with enterprise customers, researchers, product managers, and engineers. Familiarity with responsible AI, privacy-preserving analysis, data governance, and customer-data handling.

More Data jobs

Related openings