Robotics Research Infrastructure
Responsibilities Own the technical approach and ongoing reliability of a major system supporting robot data collection and experimentation.
Build and maintain robot-side and backend services, data pipelines, and tools that reduce operator effort and improve experiment throughput.
Make system health, data completeness, sensor timing, and experiment progress observable; turn failures into actionable diagnostics.
Improve deployment, configuration, testing, recovery, and rollback so changes and operations are repeatable.
Diagnose incidents across device interfaces, Linux, networking, concurrency, and storage; deliver durable fixes and verify their effect.
Prioritize work within team goals using evidence from researchers and operators, and communicate scope, risks, and tradeoffs.
Improve engineering tools and runbooks, contribute useful technical reviews, and help colleagues operate and extend the system.
Requirements Evidence of independently owning a substantial software system through design, deployment, operation, and difficult debugging.
Strong Python software engineering, including maintainable code, automated tests, concurrency, and service interfaces.
Practical Linux and networking knowledge, and experience deploying and operating services.
Understanding of distributed-system failure modes, data integrity, and safe recovery from partial failures.
Experience building observability and using it to diagnose incidents and improve reliability.
Ability to work across unfamiliar system boundaries, define an approach to ambiguous problems, and collaborate with people who use the system daily.
Nice to have Robotics, device fleets, teleoperation, edge computing, or synchronized sensor pipelines.
C++ or Rust, containers, cloud infrastructure, infrastructure as code, or ML deployment.
Operator tools and experiment platforms; effective use and critical review of AI-assisted engineering.
What Success Looks Like
Establish useful health and data-quality checks, operating documentation, and a prioritized plan for the agreed system.
Resolve a recurring failure with a measured improvement in reliability or useful data collection.
Make deployment and recovery repeatable and enable researchers and operators to diagnose routine issues