A day in the life
- Check overnight eval runs and monitoring dashboards for new failures
- Pick a weak feature, collect failing examples, and label them
- Build or tune an eval, then run it against two model versions
- Review scores with researchers and agree what counts as good enough
- Wire a passing eval into the release pipeline as a quality gate
- Write up findings and share which prompts or models to ship
Tools you will use