← Back to jobs

GenAI Security Evaluation Engineer

Location
Remote
Work type
Full Time
Posted
2026-10-06

Job description

About the Role

We’re looking for experienced security engineers to evaluate how effectively static analysis tools detect vulnerabilities in GenAI applications.

You’ll build small, runnable agent and RAG codebases containing realistic examples of Sensitive Information Disclosure and Excessive Agency, along with fixed and near-miss versions. You’ll then trace, annotate, test, and explain each finding.

What You’ll Do
Build runnable agent/RAG repositories with code-reachable Sensitive Information Disclosure or Excessive Agency vulnerabilities across tool calling, memory, and MCP.
Create vulnerable, fixed, and hard-negative variants with minimal security-relevant differences.
Trace and annotate assets, data/action paths, controls, root causes, severity, and residual risk.
Define authorization contexts and write deterministic tests validating vulnerable, fixed, and negative behavior.
Recommend security controls and participate in calibration and peer review.

What We’re Looking For
5+ years in application/product security or security-focused software engineering, including secure code review.
Experience with source-to-sink analysis, taint analysis, SAST, CodeQL, or Semgrep.
Hands-on experience building LLM agents or RAG systems using frameworks such as LangChain, LlamaIndex, OpenAI/Anthropic SDKs, or MCP.
Strong authorization knowledge, including actors, trust boundaries, tenants, OAuth, IAM, identity, permitted data/actions, purposes, recipients, and document-level access control.
Production coding experience in Python and/or TypeScript, with strong judgment in distinguishing genuine SID/EA findings from non-findings.
Experience with OWASP LLM security risks, MCP, threat modeling, or security evaluation is a plus.

Original source