← All jobs
Sonatus logo

Senior Manager, Engineering - AI Validation

SonatusWebsite ↗Sunnyvale, CAFull-time$220,000—$265,000 USD

Published 4 hours ago

The Opportunity: Sonatus is looking for an experienced Senior Engineering Manager to build and lead our AI Validation function — the team responsible for how we test, evaluate, and govern the AI models and agentic capabilities embedded in our software-defined vehicle and cloud platforms, as well as our cloud-only AI and LLM-based products. You will deeply understand how AI is developed and deployed across Sonatus’s embedded, cloud, and LLM/RAG-driven environments, identify gaps and friction in current validation practices, and turn those insights into scalable test strategy, evaluation frameworks, and governance mechanisms that let Sonatus ship trustworthy AI-driven features at automotive scale. You will lead and grow a high-performing AI validation team and partner closely with engineering, product, and safety stakeholders to make AI quality and safety a competitive advantage. Role and Responsibility: • Define and drive Sonatus's AI validation strategy across embedded, in-vehicle, cloud-connected, and cloud-native AI systems, identifying gaps in model development, testing, deployment, and governance. • Lead, hire, mentor, and grow a high-performing AI Validation organization, establishing scalable engineering processes, technical direction, and execution excellence. • Own the end-to-end validation strategy for AI/ML models, LLMs, RAG pipelines, and agentic AI workflows—from data pipelines and model training through cloud services and in-vehicle deployment. • Architect and operationalize scalable evaluation frameworks and benchmarking platforms for AI systems, including multi-step agentic workflows, using deterministic metrics, LLM-as-a-Judge methodologies, automated regression testing, and production feedback loops. • Design and maintain evaluation harnesses for RAG and agentic systems that measure retrieval quality, grounding, citation accuracy, factual consistency, context relevance, safety, latency, reliability, and execution correctness. • Evaluate, integrate, and optimize open-source and commercial AI validation technologies, driving build-versus-buy decisions for Sonatus's AI quality platform. • Establish AI governance, Responsible AI practices, model lineage, safety guardrails, and compliance processes appropriate for automotive safety-critical systems and enterprise AI products. • Act as the quality gatekeeper for AI-enabled releases, partnering with engineering, product, safety, and OEM stakeholders to identify risks, define release criteria, and ensure production readiness. • Collaborate across engineering teams to define validation strategies for emerging AI capabilities, rapidly prototype new evaluation approaches, and standardize successful practices into reusable frameworks. • Drive continuous improvement by tracking industry advances in AI evaluation, agentic AI, LLM validation, and RAG systems, translating them into scalable validation capabilities across Sonatus. Qualifications: • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field required (MS preferred). • 10+ years of experience in software or systems engineering—including embedded, cloud, networking, security, or automotive domains—with 3+ years leading high-performing engineering or QA organizations. • Hands-on experience developing, deploying, testing, or operating AI/ML systems, with strong expertise in modern ML workflows, neural networks, and MLOps. • Deep understanding of LLMs, RAG architectures, vector databases, embeddings, retrieval optimization, and agentic AI frameworks such as LangGraph or equivalent orchestration platforms. • Proven experience designing and implementing scalable evaluation frameworks for AI systems, including multi-step agentic workflows, regression testing, benchmarking, and automated quality scoring. • Strong expertise with hybrid evaluation methodologies combining deterministic validation (citation grounding, structural validation, exact matching) and probabilistic LLM-as-a-Judge techniques (faithfulness, answer relevance, context precision, task completion). • Practical experience with RAG evaluation frameworks such as RAGAS, including evaluation tuning, embedding optimization, retrieval quality improvement, and production-scale LLM evaluation pipelines. • Experience validating hallucination, grounding, citation accuracy, bias, fairness, toxicity, and factual consistency in production LLM applications. • Experience designing systems that verify external knowledge claims and ensure responses are grounded in traceable citations and trusted data sources. • Strong experience testing cloud-native platforms and cloud-managed embedded products, including end-to-end system validation. • Experience establishing AI governance, safety, compliance, and Responsible AI practices for enterprise or safety-critical systems. • Proficiency in Python, Linux, shell scripting, modern test frameworks (PyTest, Playwright, Behave), and engineering productivity tools such as Jenkins and JIRA. Ways to Stand Out: • Experience validating AI or agentic systems in safety-critical or regulated industries (automotive, aerospace, medical). • Track record building and scaling an AI test/evaluation platform or developer experience used by multiple teams (frameworks, reusable components, reference implementations). • Demonstrated wins moving AI testing practices from ad hoc to standardized, organization-wide adoption, with measurable impact on cycle time, quality, or reliability. • Experience implementing enterprise-grade AI governance (auditability, monitoring, policy enforcement) in production systems. • Deep experience evaluating LLM and RAG systems at scale, including agentic workflows, RAGAS-based evaluation, citation verification, hallucination detection, groundedness, faithfulness, answer relevance, tool/task correctness, and automated regression testing across offline and online feedback loops. Sunnyvale HQ Benefits & Perks Offered: • Health care plan (Medical, Dental & Vision) • Flexible and Dependent Care Expense program • Retirement plan (401k) • Life Insurance (Basic, Voluntary & AD&D) • Unlimited paid time off per year, 14+ paid holidays • Hybrid office work arrangement • Complimentary lunches, snacks, and beverages during on-site working days • Wellness benefit allowance • Phone & Internet reimbursement • Computer Accessory Allowance
Apply for this role
Share:LinkedInXThreads