Senior QA Engineer – AI

Delta Exchange

Description

Work Policy
Fully Remote (India)

Core Tech Stack
Python, evaluation harnesses (DeepEval/RAGAS/Promptfoo), observability (Opik/Langfuse/LangSmith)

Experience Level
4-6 years QA/SDET experience, 1+ year owning AI product quality

Job Type
Full-Time

Delta Exchange, one of the world’s most active cryptocurrency derivatives exchanges with roughly $10 billion in daily traded volume, is hiring a Senior QA Engineer – AI to own the quality of its three live AI products: a customer support chatbot, an API Copilot that generates and executes trading scripts, and an MCP server. This role is fundamentally different from typical “AI-assisted testing” roles — here, the system under test is itself an AI, and the core challenge is defining what “correct” means for a system that gives different outputs on every run.

About Delta Exchange

Delta is reimagining and rebuilding the financial system, offering the widest range of crypto derivative products and serving traders globally since 2018. The founding team includes IIT and ISB graduates with backgrounds at Citibank, UBS, and GIC, and the company is backed by leading crypto funds including Sino Global Capital, CoinFund, and Gumi Cryptos.

What You’ll Do

  • Build and maintain golden datasets for AI products, mined from production or generated, and verified against a documented source of truth.
  • Design layered scoring for AI outputs — deterministic rules where behavior can be directly asserted, and model-graded evaluation where it can’t.
  • Run evaluation cycles end to end: schedule and execute against release candidates, compare against prior baselines, and triage failures into product defects, incorrect expectations, or platform issues.
  • Own and operate the release gate that decides whether a prompt or model change gets approved for production.
  • Identify failure modes ahead of customers — hallucination, incorrect tool selection, lost conversational context, or PII exposure.
  • Investigate production traces to distinguish retrieval failures from generation failures from tool failures.
  • Turn findings into actionable engineering evidence and permanent regression coverage.
  • Define and report quality metrics for AI surfaces, and drive improvement against them.

What You’ll Need

  • At least 1 year owning quality for AI products as a lead or primary owner — chatbots, code/content generation, or agentic systems — with direct hands-on work with prompts, tool schemas, and model behavior.
  • Demonstrated hands-on experience building or operating an evaluation harness for an AI system (in-house or via DeepEval, RAGAS, Promptfoo, or similar).
  • 4-6 years of QA/SDET experience.
  • Working proficiency in Python, since the evaluation platform and tooling are Python-based.
  • Ability to analyze traces and tool calls using observability tools like Opik, Langfuse, or LangSmith to find root causes.
  • Strong API testing experience, since most of the surface under test is API-level.
  • Enough engineering ability to build your own tooling and automation.
  • Clear written communication — findings need to be documented so engineering can act on them directly.

Nice to Have

  • Experience with trading or exchange platforms, including futures, options, crypto derivatives, margin, or settlement flows.

To apply for this job please visit jobs.workable.com.