Available for new projects

Nikita Kalachev QA & LLM Evaluation Engineer

Portrait of Nikita Kalachev, QA and LLM Evaluation Engineer

I find where software quietly fails — now in LLM outputs and the products built on them.

Remote · Independent contractor

7 years
QA on business-critical systems
1M+ users
on the largest platform I tested
4 providers
run against each other in CrossCheck
~1,300 tests
across 103 test files in CrossCheck

Summary

I spent seven years finding bugs in software companies could not afford to break: checkout systems, payment flows, medical records. Four and a half of those years were at EPAM Systems, on platforms used by more than a million people.

For the past year I have been doing the same work on AI. I built CrossCheck, a tool that sends one question to several AI models and shows where they disagree. Most tools give you a single answer. This one shows you which parts of it to distrust.

I now take contract work in LLM evaluation and QA for AI products.

What I Do

01

LLM output evaluation

I check model outputs for accuracy, invented facts, whether the model followed the instruction, and whether the reasoning holds. I write the criteria down so a second person grading the same output gets a similar result.

02

Evaluation methodology and rubrics

Teams usually know one answer is worse but cannot say why. I turn that into a written scale with weights, so the score can be checked and argued with.

03

Adversarial testing and red-teaming

I build test sets meant to break the model: prompt injection, edge cases, and the failures that only show up under pressure.

04

QA for AI products

Release testing for teams shipping AI features. Regression on what the model says, not just on the interface around it.

Most engagements start with a small paid trial task — the low-risk way to see the work before committing. Start a trial task →

A tool I built

CrossCheck — to catch AI disagreement

CrossCheck sends the same question to several AI models, has them argue with each other, and reports where they disagree. Most tools hand you one confident answer. This one shows you the parts the models could not agree on, so you know what to check.

It runs in three stages. Each model answers on its own. Then they critique each other, two rounds in the standard setting. Then everything is combined into one report that keeps the disagreements visible.

The score

Every disagreement is weighted by how serious it is: 20 points off for a critical conflict, 10 for a moderate one, 5 for a minor one. It is plain arithmetic, and you can read the function and check the math. No model grades another model.

What it does not do

There is no automated benchmark for the verification itself, and I do not pretend otherwise. I tested the behaviour by running cases by hand. The scoring math is covered by unit tests. One part of the system, the PII scrubber, has a real measured benchmark against a synthetic test set.

Under the hood

Four providers available to users: Anthropic, OpenAI, Google, Mistral. Up to eleven models depending on which keys you add. A fifth adapter, AWS Bedrock, is wired up but internal. Roughly 1,300 automated tests, and about as much test code as production code. Next.js, TypeScript, Supabase. Public beta since May 2026, no revenue yet.

How I Think About Evaluation

I write about verification and where AI reliability claims break down.

Experience

Founder & Builder

2026 – Present · Remote
CrossCheck AI (Platilus LLC)

Built a multi-model LLM verification system on my own, from the first problem sketch to public beta.

  • Designed the verification method, the scoring model and the architecture, and made the product calls: bring your own key, so the user's content stays theirs; show disagreement instead of averaging it into one confident answer.
  • Wrote the rules the AI coding agents had to work under. Every step has to prove itself against something outside the model — the compiler, the tests, or a written architecture contract. A plan without a quote from real code gets rejected. Three failed attempts on the same problem and the work stops for a human decision.
  • Audited my own product's claims against the codebase and removed the ones the code did not support: a check that existed only in the spec, a claim about repeatable runs that ignored uncontrolled temperature upstream, and a provider that was never integrated.
  • Kept the suite at 1,313 tests across 103 files — about as much test code as production code — because generated code needs more checking, not less. There is no end-to-end suite and no integration coverage, and I say that plainly rather than implying more.
  • Built the cost and reliability limits: stop before calling a model if the estimate is too high, a ceiling per session, a daily cap per user. One provider failing does not bring down a verification run.

QA Engineer

Sep 2021 – Feb 2026
EPAM Systems · Remote

Four and a half years on enterprise projects for large retail, healthcare and fintech clients.

  • Loyalty platform for over a million customers across 12 European markets. Found 50+ critical bugs before release, validated 200+ API endpoints.
  • Checkout for a global retailer handling 10,000+ orders a day. Integration testing and UAT, payment gateways in 8 countries under PCI rules. No critical bugs through Black Friday.
  • Healthcare admin portal for 500,000 members under HIPAA. Wrote 100+ UI automation scripts, raised regression coverage by 40%.
  • Mobile banking app with AI budgeting features, iOS and Android. Payment processing and data sync.

QA Team Lead

May 2021 – Sep 2021
Healthcare Technology Company

Led a team of four on a dental clinic CRM: booking, billing, inventory.

QA Project Lead

May 2020 – Apr 2021
Web Development Agency

Full-cycle QA on 8+ projects in retail, telecom, finance and healthcare, all shipped with no critical defects after launch. Cross-browser and security testing.

Skills

QA
manual functional regression API integration UAT cross-browser mobile exploratory testing test design root-cause analysis
AI and evaluation
LLM output evaluation hallucination detection instruction-following reasoning quality adversarial testing red-teaming prompt injection multi-model comparison rubric design prompt engineering
Tools
Postman Swagger Charles Proxy Jira Azure DevOps BrowserStack Vividus Zephyr Scale SQL Git Docker TypeScript Next.js Supabase
Domains
e-commerce fintech healthcare payment gateways (PCI) HIPAA
Languages
English (B2) Russian (native)

FAQ

What kind of evaluation work do you take on?

LLM output evaluation, evaluation methodology and rubric design, adversarial test sets and red-teaming, and QA for teams shipping AI features. I am strongest where the judgment is about correctness, reasoning, instruction-following, and safety of outputs. I do not take work that depends on line-by-line code judgment in languages I do not work in, or on native-level English style judgment.

How did you build CrossCheck?

I directed AI coding agents instead of writing the code by hand. I designed the method, the scoring model and the architecture, set the rules the agents had to follow, and checked what came back. That is why the project has around 1,300 tests: generated code needs more checking, not less. It is also how I ended up in AI evaluation — reviewing AI output is what I do every day.

Are you available?

Yes. I take contract work through my company, Platilus LLC, with a contract and invoices. CrossCheck is a public beta with no revenue, so it is not a company I am hiring for. A small paid trial task is usually the easiest way to start. I work in English (B2) and Russian (native). Write to [email protected].

Available for evaluation & QA contracts.

A small paid trial task is usually the easiest way to start — contract and invoices through Platilus LLC.