AI Systems Evaluation & Reliability Internship
- Location
- Remote
- Work type
- Part Time · Remote
- Posted
- 2026-08-05
Job description
What You Will Work On
Run and analyze memory/recall benchmarks.
Add benchmark controls and ablations, such as no-context baselines, retrieval-vs-reader failure analysis, and regression checks.
QA agent-generated PRs by running tests, reading logs, checking edge cases, and validating behavior.
Write small Python, Rust, and SQL tools for benchmark analysis, smoke tests, and CI validation.
Maintain concise runbooks so benchmark and development workflows are repeatable.
Help summarize results clearly for the team, including what passed, what failed, and what the result actually means.
Our Ideal Candidate
Strong CS fundamentals and fast learning ability.
Comfortable with terminal workflows, Git/GitHub, and CI logs.
Python for benchmark/eval scripting.
Rust interest or experience, since the product codebase is Rust-first.
Basic SQL/Postgres familiarity.
Good testing and debugging instincts.
Familiarity with or strong interest in LLM/RAG concepts: embeddings, vector search, evals, model-as-judge, context windows.
Clear written communication and high attention to detail.
High agency: able to take an ambiguous validation task and make steady progress.
Nice To Have
Rust project experience.
Docker or CI experience.
Prior ML/AI project.
Basic statistics: accuracy, confidence intervals, sampling, ablations.
Ideal Project For July-August
Help build the validation layer around Fabric: benchmark runbooks, no-context controls, regression smoke tests, PR QA checklists, failure analysis reports, and lightweight tools that let us ship agent-generated code faster without losing correctness.