Testing LLM-based applications

avatar
Ludovico Besana
26 Sept 2025
  • Share

From the stage of We Make Future 2025 to CI/CD pipelines: how to build quality strategies for generative AI applications, even in a non-deterministic context.

At We Make Future 2025, I presented a challenge every team building with AI now faces: how do we test something that never behaves the same way twice?

Representing Nearform, I presented a topic that is now central to anyone working in software quality: how to test applications based on LLMs (Large Language Models) in a non-deterministic context.

Why does this matter? Because AI apps don’t work like traditional software: the same prompt can produce different answers. This breaks the very concept of deterministic testing, but it doesn’t mean we should give up on quality.

In this article I’ll share the main strategies, tools, and takeaways from my talk, along with the most common questions I received on stage.

Ludovico Besana, Senior Test Engineer at Nearform, on stage at We Make Future 2025 during the talk 'The AI Bug: How to Test Applications with Unpredictable Outputs

What’s the problem with testing AI?

  • Two executions of the same prompt can generate different responses.
  • A single “correct” answer doesn’t always exist.

The classic test oracle doesn’t hold up. That’s why we need a multi-layered approach that combines automation, new metrics, and human oversight.

Diagram showing the relationship between different testing levels: Unit Testing, Functional Testing, Responsibility Testing, and Performance Testing, all within the context of Regression Testing.

The 5 levels of LLM testing

During the talk I proposed a 5-level model to progressively address complexity:

1. Unit Test

Prompt with a clear and unambiguous answer.

Example: “What is the capital of France?”

Metric: Exact Match

2. Integration Test

LLM + API + RAG + plugins.

Metric: End-to-end success

3. Functional & Performance Test

Complex tasks (translations, summaries, code generation).

Metrics: BLEU, ROUGE, Cosine Similarity, TPS (tokens/sec), latency.

4. Regression & Stress Test

Versioned datasets, multiple re-runs, adversarial prompts.

Metric: % errors, delta score

5. Human-in-the-Loop

Human evaluation of tone, coherence, and clarity.

Metric: Checklist or qualitative scale

Hallucination is a bug

One of the most critical problems in LLM systems is hallucination: the model invents data that sounds plausible but is false.

Examples:

  • Code using deprecated libraries
  • Nonexistent citations
  • Fabricated numbers or dates

In regulated fields (finance, healthcare, legal), this can have massive consequences.

We shouldn’t treat it as a “creative feature,” but as a bug to catch.

The limits of AI testing

Even with best practices, some challenges cannot be eliminated:

LimitPractical Impact
Non-determinismDifferent outputs from the same prompt
Infinite combinationsImpossible to cover all cases
Fragile golden referenceOracle may be inaccurate
Bias & safetyOften require human review
Real costsEvery test = token consumption

Tools I'm using

Here are the main tools I use for AI testing:

I’m also testing Llama 3 locally with DeepEval to avoid API token costs.

LLM testing in CI/CD

AI testing can fit into pipelines.

Useful practices include:

  • Integrating tests on critical prompts at every release
  • Versioning datasets and internal benchmarks
  • Monitoring regressions with dedicated dashboards

This way you can track improvements or regressions over time.

Questions from the audience

During the Q&A session at WMF, several practical questions came up. For example:

  • How do you test email marketing or legal content?

    → Define a range of valid answers instead of a single golden one.

  • Can we remove humans from the loop?

    → Partially yes, with tools like DeepEval that run LLM vs LLM, but human judgment remains key.

  • Are there reliable open-source models?

    → Yes. Some general-purpose models (Mistral, O1) are solid. Local testing + prompt adaptation is the key.

Conclusion

Testing AI isn’t about proving it always works—it’s about knowing where it breaks, how to measure it, and how to keep improving.

The key is balance:

  • Automate the checks that are repeatable and scalable.
  • Evaluate the outputs with meaningful metrics and human judgment.
  • Adapt your strategy as models, data, and risks evolve.

Quality in the LLM era means embracing uncertainty without sacrificing trust. At Nearform, we help teams do exactly that—building AI systems that are innovative and reliable, so they can move fast without losing confidence. Reach out to us here.

Ludovico Besana and other speakers on stage at We Make Future 2025 in front of a screen displaying 'Thank you :)'. A photographer is taking their picture while the audience watches.

But wait - there's more.

Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.

Insights

Perspectives on AI in engineering, product development, and strategy, for enterprise executives.

Community

Deep dives and tutorials by engineers, for engineers.

Insight, imagination and expertly engineered solutions to accelerate and sustain progress.