Job Description
About the Role
We are seeking a Lead AI/ML Engineer to architect, build, and operationalize a full-lifecycle, automated evaluation harness for non-deterministic Generative AI agents deployed across our Cloud GTM space (e.g., deep research, conversational analytics, pipeline forecasting).
In this role, you will transition our AI evaluations from manual human reviews to high-throughput automated pipelines backed by LLM-as-a-Judge and targeted human audits. You will work at the intersection of enterprise software engineering, distributed computing, and cutting-edge LLM evaluation frameworks to ensure data integrity, eliminate regressions, and scale our GTM agent capabilities.
Key Responsibilities 1. Foundations & Local Sandbox Development Build robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos, phrasing, syntax) for robust testing. Construct local developer sandbox environments using the Google ADK fram...