A reliability metric the client had never measured
MTBF (mean time between failures) for vehicle telemetry services: designed, built and shipped in under three months, AI-first.
- Problem
- How long does the car's telemetry run without failures? The question went to 5 manual-testing teams spread over 3 offices. Nobody had tested this before. Then it was handed to me.
- Approach
- I sat with the teams to learn their tools and pain points, and turned them into specs for agentic coding. Then I built the whole stack: scheduled Robot Framework runs driving SOME/IP, CAN and offboard APIs, a PostgreSQL trace store, GitLab CI, and the network paths into each office, negotiated with each office's IT department.
- Outcome
- In production 24/7 since month three:
- ~430 tests a day
- every test run is traceable to every command sent and its output
- KPIs and bench utilisation statistics are used by the head of the client's European testing
- scalable to dozens of benches
- Robot Framework
- FastAPI
- PostgreSQL
- Grafana
- APScheduler
- Docker
- GitLab CI
- SOME/IP
- CAN
- ADB
- VS Code Copilot
- Pi agent

