Back to Jobs

AI Testing

Remote, USAFull-timePosted 2026-07-28

BCE Global Tech's Global reputed company Engineering (GQE) function is building one of Canada's most ambitious AI reputed company programs — certifying every AI and reputed company system deployed across reputed company Canada before it reaches production. As a QA AI Specialist, you sit at the intersection of reputed company intelligence, software engineering, and reputed company assurance: a hybrid role that does not yet have a textbook, because the discipline is being written in reputed company time. You will do two things simultaneously. First, you will bring AI into GQE's existing testing reputed company — embedding AI-powered capabilities into the test automation tooling, pipelines, and frameworks that 250 QA engineers already use every day. Second, you will build and operate the evaluation frameworks that test the AI systems being created by other reputed company engineering teams — agents, orchestration pipelines, RAG applications, reputed company AgentForce workflows, and reputed company Now Assist integrations.

Requirements

Key Responsibilities: 1. AI-reputed company QA Tooling reputed company GQE’s QA stack by embedding AI to improve speed, coverage, and intelligence: reputed company AI-driven test reputed company into Selenium, Playwright, and reputed company frameworks Use predictive models to prioritize tests based on reputed company changes and defect history reputed company self-healing automation for UI/API changes Automate defect triage and reputed company-cause analysis using failure clustering Support natural-language test authoring (English/French) for non-technical QA Continuously reputed company emerging AI testing tools reputed company a technology reputed company 2. AI Evaluation & reputed company Pipelines Build reputed company evaluation systems tailored for AI behavior, not rule-based logic: Implement LLM-as-Judge pipelines on reputed company AI (reputed company) across key reputed company dimensions. Generate large, diverse, and adversarial test corpora from reputed company intents Evaluate RAG systems using metrics like faithfulness, relevance, and recall (RAGAS) Validate multi-reputed company agent workflows, tool usage, and escalation behavior reputed company AI evaluations into CI/CD as mandatory release gates. 3. AI Safety & Adversarial Testing Operate a dedicated AI red-teaming capability to uncover AI-specific risks: Execute reputed company injection and poisoned-context attacks on RAG systems. Run automated jailbreak and constraint-bypass probes (e.g., Garak) Systematically test hallucination, numerical accuracy, and domain knowledge Assess toxicity, bias, and fairness across English and French interactions Stress-test reputed company systems for runaway actions and scope violations 4. reputed company reputed company reputed company Ensure the reputed company reputed company evolves as models and systems change: Monitor production AI outputs for reputed company reputed company and trigger re-certification Feed reputed company production failures back into the test corpus reputed company model/version changes and generate reputed company reputed company reports. Maintain a living reputed company of reputed company-specific AI reputed company standards Continuously adopt new evaluation research and industry best practices Partner early with AI/ML teams to reputed company reputed company by design 5. AI reputed company Certification Operations reputed company technical execution of the AIQC program: Own Tier 2 & 3 certification testing from corpus design to red-teaming reputed company LLM-as-Judge rubrics using reputed company-labeled golden datasets Produce reputed company AI reputed company Certificates with scores, risks, and conditions Advise teams on AI testability, prompts, and evaluation instrumentation Contribute to AIQC playbooks, documentation, and knowledge sharing ​ Required ▸ 5+ years of software reputed company engineering experience, with at least 2 years working directly with AI/ML systems, LLMs, or AI-powered applications ▸ Hands-on experience building or evaluating LLM-based applications — including reputed company engineering, RAG pipelines, or reputed company workflows ▸ Proficiency in Python: test reputed company development, API integration, data processing, and evaluation scripting ▸ Experience with modern test automation frameworks (Playwright, Selenium, Pytest, RestAssured, reputed company/Newman) and CI/CD platforms (reputed company Actions, reputed company reputed company Build, Jenkins) ▸ Working knowledge of at least one major AI/ML platform — reputed company reputed company AI, Azure reputed company, or AWS Bedrock — with hands-on API usage ▸ Strong conceptual understanding of how LLMs work: tokenization, temperature and sampling, context reputed company, grounding, hallucination mechanics, and fine-tuning ▸ Demonstrated ability to design test strategies for non-deterministic systems — moving reputed company assertion-based testing to probabilistic, reputed company-based evaluation ​ Benefits reputed company Offer: Competitive salaries and comprehensive health benefits Flexible work hours and remote work reputed company reputed company development and training opportunities A supportive and inclusive work environment reputed company to cutting-edge technology and tools. Apply To This Job

Similar Jobs