Senior Manager, Engineering - AI Validation
Published 4 hours ago
The Opportunity:
Sonatus is looking for an experienced Senior Engineering Manager to build and lead our AI Validation function — the team responsible for how we test, evaluate, and govern the AI models and agentic capabilities embedded in our software-defined vehicle and cloud platforms, as well as our cloud-only AI and LLM-based products. You will deeply understand how AI is developed and deployed across Sonatus’s embedded, cloud, and LLM/RAG-driven environments, identify gaps and friction in current validation practices, and turn those insights into scalable test strategy, evaluation frameworks, and governance mechanisms that let Sonatus ship trustworthy AI-driven features at automotive scale. You will lead and grow a high-performing AI validation team and partner closely with engineering, product, and safety stakeholders to make AI quality and safety a competitive advantage.
Role and Responsibility:
• Define and drive Sonatus's AI validation strategy across embedded, in-vehicle, cloud-connected, and cloud-native AI systems, identifying gaps in model development, testing, deployment, and governance.
• Lead, hire, mentor, and grow a high-performing AI Validation organization, establishing scalable engineering processes, technical direction, and execution excellence.
• Own the end-to-end validation strategy for AI/ML models, LLMs, RAG pipelines, and agentic AI workflows—from data pipelines and model training through cloud services and in-vehicle deployment.
• Architect and operationalize scalable evaluation frameworks and benchmarking platforms for AI systems, including multi-step agentic workflows, using deterministic metrics, LLM-as-a-Judge methodologies, automated regression testing, and production feedback loops.
• Design and maintain evaluation harnesses for RAG and agentic systems that measure retrieval quality, grounding, citation accuracy, factual consistency, context relevance, safety, latency, reliability, and execution correctness.
• Evaluate, integrate, and optimize open-source and commercial AI validation technologies, driving build-versus-buy decisions for Sonatus's AI quality platform.
• Establish AI governance, Responsible AI practices, model lineage, safety guardrails, and compliance processes appropriate for automotive safety-critical systems and enterprise AI products.
• Act as the quality gatekeeper for AI-enabled releases, partnering with engineering, product, safety, and OEM stakeholders to identify risks, define release criteria, and ensure production readiness.
• Collaborate across engineering teams to define validation strategies for emerging AI capabilities, rapidly prototype new evaluation approaches, and standardize successful practices into reusable frameworks.
• Drive continuous improvement by tracking industry advances in AI evaluation, agentic AI, LLM validation, and RAG systems, translating them into scalable validation capabilities across Sonatus.
Qualifications:
• Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field required (MS preferred).
• 10+ years of experience in software or systems engineering—including embedded, cloud, networking, security, or automotive domains—with 3+ years leading high-performing engineering or QA organizations.
• Hands-on experience developing, deploying, testing, or operating AI/ML systems, with strong expertise in modern ML workflows, neural networks, and MLOps.
• Deep understanding of LLMs, RAG architectures, vector databases, embeddings, retrieval optimization, and agentic AI frameworks such as LangGraph or equivalent orchestration platforms.
• Proven experience designing and implementing scalable evaluation frameworks for AI systems, including multi-step agentic workflows, regression testing, benchmarking, and automated quality scoring.
• Strong expertise with hybrid evaluation methodologies combining deterministic validation (citation grounding, structural validation, exact matching) and probabilistic LLM-as-a-Judge techniques (faithfulness, answer relevance, context precision, task completion).
• Practical experience with RAG evaluation frameworks such as RAGAS, including evaluation tuning, embedding optimization, retrieval quality improvement, and production-scale LLM evaluation pipelines.
• Experience validating hallucination, grounding, citation accuracy, bias, fairness, toxicity, and factual consistency in production LLM applications.
• Experience designing systems that verify external knowledge claims and ensure responses are grounded in traceable citations and trusted data sources.
• Strong experience testing cloud-native platforms and cloud-managed embedded products, including end-to-end system validation.
• Experience establishing AI governance, safety, compliance, and Responsible AI practices for enterprise or safety-critical systems.
• Proficiency in Python, Linux, shell scripting, modern test frameworks (PyTest, Playwright, Behave), and engineering productivity tools such as Jenkins and JIRA.
Ways to Stand Out:
• Experience validating AI or agentic systems in safety-critical or regulated industries (automotive, aerospace, medical).
• Track record building and scaling an AI test/evaluation platform or developer experience used by multiple teams (frameworks, reusable components, reference implementations).
• Demonstrated wins moving AI testing practices from ad hoc to standardized, organization-wide adoption, with measurable impact on cycle time, quality, or reliability.
• Experience implementing enterprise-grade AI governance (auditability, monitoring, policy enforcement) in production systems.
• Deep experience evaluating LLM and RAG systems at scale, including agentic workflows, RAGAS-based evaluation, citation verification, hallucination detection, groundedness, faithfulness, answer relevance, tool/task correctness, and automated regression testing across offline and online feedback loops.
Sunnyvale HQ Benefits & Perks Offered:
• Health care plan (Medical, Dental & Vision)
• Flexible and Dependent Care Expense program
• Retirement plan (401k)
• Life Insurance (Basic, Voluntary & AD&D)
• Unlimited paid time off per year, 14+ paid holidays
• Hybrid office work arrangement
• Complimentary lunches, snacks, and beverages during on-site working days
• Wellness benefit allowance
• Phone & Internet reimbursement
• Computer Accessory Allowance