ABOUT THE ROLE
We're a well-funded, early-stage startup building AI agent technology that turns mechanical engineers' intent into reliable, cost-efficient multi-step workflows across complex desktop engineering tools. Our product sits at the intersection of applied agentic AI, enterprise software automation, and deep domain expertise in CAD, simulation, and PLM.
As Staff Engineer — Agentic AI, you will own the core agent intelligence layer of our product. This is a high-leverage, high-visibility role: you'll report directly to the CTO, serve as technical lead for a small team of AI engineers, a user researcher, and domain expert contractors, and be the person most responsible for whether the agent actually delivers value to Fortune 100 engineering teams.
This role is on-site in San Francisco, CA. Visa sponsorship is not available.
WHAT YOU'LL DO
- Drive agent task success rate. Own the metric that matters most — can the agent actually complete the workflows engineers need? Define the eval framework, establish baselines, and systematically improve performance.
- Set and enforce token budgets per workflow. Define per-task token budgets, track cost per completed workflow, and make the agent commercially viable — not just technically impressive.
- Build rigorous evaluation infrastructure. Design benchmarks grounded in real user stories, not synthetic tasks. Think SWE-bench-level rigor applied to engineering workflows — reproducible, adversarial, and tied to actual customer value.
- Lead user story mapping and validation. Work directly with the user researcher and domain expert contractors to interview engineers, document their workflows in detail, and validate that what you're building against reflects reality — not assumptions.
- Translate user stories into evals. Every validated user story becomes a test case. Close the loop between user research and agent benchmarking.
- Expand workflow coverage. Systematically increase the percentage of steps in top user stories that the agent can handle end-to-end, prioritizing by customer value and technical feasibility.
- Own the agent architecture. Make foundational decisions around tool-calling strategies, state management across multi-step workflows, error recovery, model routing, and context management.
- Lead a team as a player-coach. Set technical direction, review architecture decisions, unblock the team, raise the engineering bar — and still write production code.
- Collaborate cross-functionally with integrations, product, and enterprise customers during POCs to align agent behavior with real-world usage.
WHAT WE'RE LOOKING FOR
Dealbreakers (must-haves):
- 7+ years in software engineering, including at least 2 years building agentic LLM-based systems that call tools, manage multi-step workflows, handle failures, and operate under cost constraints.
- Deep experience with LLM application architecture: model selection, context window management, retrieval strategies, tool-calling frameworks, and orchestration.
- Strong evaluation and benchmarking instincts for agentic systems — task completion, cost efficiency, failure mode analysis; familiarity with benchmarks such as SWE-bench, GAIA, or τ-bench.
Required:
- Proven track record of shipping AI systems with measurable outcomes (e.g., agent task success rate, cost efficiency) — not just demos.
- Strong Python skills and working knowledge of the LLM tooling ecosystem: function calling, tool use APIs, tracing/observability tools (e.g., Logfire, LangSmith), and evaluation frameworks.
- Experience leading a small technical team (3–6 engineers): setting technical direction, performing code reviews, and driving architecture decisions.
Nice to Have:
- Published work or open-source contributions in agentic AI systems.
- Experience with desktop automation, COM, or programmatic control of applications (beyond web APIs).
- Background in mechanical engineering, CAD/CAE, PLM, or adjacent industries.
- Familiarity with enterprise deployment constraints — running agents on locked-down corporate workstations.
- Experience building or contributing to public benchmarks for AI agents.
COMPENSATION & BENEFITS
- Salary: $160,000 – $250,000 per year, depending on experience
- Early-stage equity in a Series A company with strong investor backing
- Direct line to the CTO and outsized impact at a critical stage of company growth
- Small, high-caliber team working on a technically deep problem in a largely untapped industry vertical
LOCATION
This is an on-site role based in San Francisco, CA. Candidates must be authorized to work in the United States — visa sponsorship is not available.