Careers / IT
AI QA Engineer (Remote)
Help define how intelligent systems are tested and trusted. Design and execute quality strategies for generative AI, LLM, RAG, agentic, and API-driven applications across functional quality, safety, reliability, and responsible AI.
Responsibilities
- Develop end-to-end AI quality strategies covering functional behavior, model outputs, APIs, data flows, integrations, and user journeys.
- Design evaluation datasets, test scenarios, prompts, expected outcomes, scoring rubrics, and acceptance thresholds for LLM, RAG, and agentic applications.
- Evaluate response accuracy, relevance, groundedness, consistency, completeness, hallucination risk, toxicity, bias, privacy, and safety.
- Test retrieval pipelines, source attribution, context handling, prompt injection defenses, fallback behavior, and human-in-the-loop workflows.
- Build and maintain automated tests using Python, Playwright, API tooling, and reusable evaluation frameworks.
- Perform exploratory, regression, negative, boundary, adversarial, and risk-based testing across model and application changes.
- Define measurable quality indicators and release criteria; analyze trends and communicate findings through clear dashboards and evidence.
- Partner with product, engineering, data science, governance, and business stakeholders to translate requirements and AI risks into testable controls.
- Document defects with reproducible evidence, isolate likely failure causes, and validate fixes across prompts, models, data, orchestration, and application code.
- Integrate AI evaluations into CI/CD pipelines and help establish quality gates that prevent unsafe or low-quality releases.
- Contribute reusable test assets, evaluation patterns, and lessons learned to strengthen AiQualTest delivery practices.
Qualifications
- 5+ years of software quality engineering or test automation experience, including hands-on testing of AI/ML, generative AI, data-intensive, or API-driven applications.
- Practical understanding of LLM behavior and common failure modes such as hallucination, weak grounding, prompt sensitivity, inconsistent output, bias, and unsafe responses.
- Experience creating structured test strategies, evaluation datasets, traceable test cases, acceptance criteria, and executive-ready quality reports.
- Strong API testing skills and proficiency with Python or another language used to automate evaluations and analyze results.
- Experience with browser or end-to-end automation such as Playwright, Cypress, or Selenium.
- Ability to test probabilistic systems where outcomes require scoring, thresholds, repeated runs, and statistical interpretation rather than exact-match assertions alone.
- Working knowledge of data quality, privacy, security, accessibility, and responsible AI considerations.
- Strong analytical thinking and the ability to distinguish model, data, prompt, retrieval, integration, and application-layer defects.
- Clear written and verbal communication for collaborating with distributed technical and non-technical teams.
- Ability to work independently in a remote environment and manage priorities across multiple evaluation workstreams.
Schedule & terms
- Weekly pay options for contractors
- Remote flexibility
- Recruiter advocacy and interview prep
Nice to have
- AiVELabs Learn → Practice → Validate ecosystem
- Hands-on AI quality labs
- Responsible AI and AI governance
- NIST AI RMF
- ISO/IEC 42001
- RAG evaluation
- CI/CD quality gates
- ISTQB certification
Familiarity with the AiVELabs Learn → Practice → Validate ecosystem is especially valuable. Experience completing role-based AI assurance learning, practicing in hands-on labs, or applying those skills to real-world AI evaluations will help you contribute quickly.
Apply or refer
Equal opportunity
AiQualTest is an Equal Opportunity Employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We encourage applications from candidates of all backgrounds for both IT and Non-IT placements.
Related openings
- Sales & Upskilling (Product + Bench)
Remote (India) · Performance-based commissions
- Pure Bench Sales
Remote (India) · Performance-based placement commissions
- SDET
Remote (US)