At We Make Future 2025, I presented a challenge every team building with AI now faces: how do we test something that never behaves the same way twice?
Representing Nearform, I presented a topic that is now central to anyone working in software quality: how to test applications based on LLMs (Large Language Models) in a non-deterministic context.
Why does this matter? Because AI apps don’t work like traditional software: the same prompt can produce different answers. This breaks the very concept of deterministic testing, but it doesn’t mean we should give up on quality.
In this article I’ll share the main strategies, tools, and takeaways from my talk, along with the most common questions I received on stage.

What’s the problem with testing AI?
- Two executions of the same prompt can generate different responses.
- A single “correct” answer doesn’t always exist.
The classic test oracle doesn’t hold up. That’s why we need a multi-layered approach that combines automation, new metrics, and human oversight.

The 5 levels of LLM testing
During the talk I proposed a 5-level model to progressively address complexity:
1. Unit Test
Prompt with a clear and unambiguous answer.
Example: “What is the capital of France?”
Metric: Exact Match
2. Integration Test
LLM + API + RAG + plugins.
Metric: End-to-end success
3. Functional & Performance Test
Complex tasks (translations, summaries, code generation).
Metrics: BLEU, ROUGE, Cosine Similarity, TPS (tokens/sec), latency.
4. Regression & Stress Test
Versioned datasets, multiple re-runs, adversarial prompts.
Metric: % errors, delta score
5. Human-in-the-Loop
Human evaluation of tone, coherence, and clarity.
Metric: Checklist or qualitative scale
Hallucination is a bug
One of the most critical problems in LLM systems is hallucination: the model invents data that sounds plausible but is false.
Examples:
- Code using deprecated libraries
- Nonexistent citations
- Fabricated numbers or dates
In regulated fields (finance, healthcare, legal), this can have massive consequences.
We shouldn’t treat it as a “creative feature,” but as a bug to catch.
The limits of AI testing
Even with best practices, some challenges cannot be eliminated:
| Limit | Practical Impact |
|---|---|
| Non-determinism | Different outputs from the same prompt |
| Infinite combinations | Impossible to cover all cases |
| Fragile golden reference | Oracle may be inaccurate |
| Bias & safety | Often require human review |
| Real costs | Every test = token consumption |
Tools I'm using
Here are the main tools I use for AI testing:
- Promptfoo → regression testing, injection, security
- DeepEval → output evaluation + comparison
- BenchLLM → continuous testing in CI/CD
- LangChain → complex test flows locally
- HuggingFace Evaluate → standard metrics (BLEU, ROUGE, METEOR)
I’m also testing Llama 3 locally with DeepEval to avoid API token costs.
LLM testing in CI/CD
AI testing can fit into pipelines.
Useful practices include:
- Integrating tests on critical prompts at every release
- Versioning datasets and internal benchmarks
- Monitoring regressions with dedicated dashboards
This way you can track improvements or regressions over time.
Questions from the audience
During the Q&A session at WMF, several practical questions came up. For example:
-
How do you test email marketing or legal content?
→ Define a range of valid answers instead of a single golden one.
-
Can we remove humans from the loop?
→ Partially yes, with tools like DeepEval that run LLM vs LLM, but human judgment remains key.
-
Are there reliable open-source models?
→ Yes. Some general-purpose models (Mistral, O1) are solid. Local testing + prompt adaptation is the key.
Conclusion
Testing AI isn’t about proving it always works—it’s about knowing where it breaks, how to measure it, and how to keep improving.
The key is balance:
- Automate the checks that are repeatable and scalable.
- Evaluate the outputs with meaningful metrics and human judgment.
- Adapt your strategy as models, data, and risks evolve.
Quality in the LLM era means embracing uncertainty without sacrificing trust. At Nearform, we help teams do exactly that—building AI systems that are innovative and reliable, so they can move fast without losing confidence. Reach out to us here.

But wait - there's more.
Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.
Insights
Perspectives on AI in engineering, product development, and strategy, for enterprise executives.
Community
Deep dives and tutorials by engineers, for engineers.
You may also like

Nearform’s LLM Playground: We created our version of the OpenAI Playground

