Sideways, not upwards: Scaling frontend performance with k6 (and why hybrid wins)

avatar
Raluca Chirila
16 Dec 2025
  • Share

In under a week, we proved that pre-production could sustain approximately 200 concurrent users with the k6 browser module, giving the team the confidence to release

Most modern teams today understand the crucial role performance testing plays in assessing the performance and scalability of a web or mobile application.

Performance testing is all about checking how a system behaves under load - how fast it responds, how much traffic it can handle, and when it starts to struggle. Most importantly, it is about understanding what users feel when the system is under load.

The challenge is how we put it into practice.

Classic protocol-level tools such as Gatling, JMeter, LoadRunner have been around for many years. They are great at testing backend APIs, simulating real-world user traffic, and giving us insights into aspects like capacity, reliability and responsiveness. However, they are less effective at conveying the actual user experience when a real browser renders the page under load on the server.

On the other hand, we have browser-based tools like Playwright or Cypress. Simply put, these are excellent at end-to-end checks and catching frontend bugs, and they can also be extended into a bit of performance testing, but they’re not built for actual distributed load on their own.

k6

Where does k6 sit in the performance testing landscape?

k6 is an open-source, code-first performance testing tool from Grafana, designed for load testing web applications and APIs. It is well known for its simplicity and integration with the modern development ecosystem. k6 uses Javascript/Typescript for writing test scripts which can be committed like any other code and gives you performance metrics out of the box. It supports performance checks and thresholds to validate the system’s behaviour and set pass / fail criteria.

More recently, the k6/browser module added support for browser automation and end-to-end web testing whilst still preserving its core k6 features. Moreover, it uses Playwright-style selectors, so we can run both protocol-level and browser-level scenarios from the same ecosystem.

k6 Operator

The k6 Kubernetes Operator takes things one step further, enabling distributed k6 tests in your cluster. This is particularly useful when a single node cannot handle the load required by your test, or simply when Kubernetes is already your preferred operations environment.

The same script you run locally can be published into your cluster and spread across multiple pods, each running a chunk of the load, without having to manually orchestrate containers and nodes.

That’s the “happy path” story:

  • design a load model,
  • write k6 scripts (HTTP + optional browser),
  • scale them out with the k6 Operator depending on your use case.

The challenge

Our story starts where that neat story breaks down.

Time wasn’t on our side before a major release. Our usual performance testing route into pre-production was blocked. Testing performance via the backend APIs wasn’t an option. Still, we needed confidence in the user experience under load.

The solution wasn’t our regular go-to. However, it proved to be practical.

We kept our north star, hybrid testing with k6: cheap, high-volume HTTP for backend targets and browser tests for the human feel. And when pre-production constraints closed the HTTP door, we had to go browser-only. So, we scaled horizontally with the k6 Operator on OpenShift (OCP) and were able to push far enough to validate capacity with browser user journeys.

Hybrid by default

Under normal circumstances, hybrid performance testing (built with k6) is the sweet spot.

Why? Because HTTP scenarios give you high-volume arrival-rate at pennies per request. Simultaneously, browser flows give you web vitals metrics, the kind that product and design actually care about. So, we use HTTP tests to simulate realistic load models, and a thin slice of browser journeys to put numbers on what users feel.

Because k6’s browser module uses Playwright-style selectors and patterns, engineers who already write Playwright E2E tests will feel at home here, as k6 uses similar selectors, but with k6’s distributed execution and performance-centric metrics.

In short, use HTTP to load the system, and the browser to tell you what the human would have felt under that load.

The pre-prod constraint (and our first face-plants)

But pre-prod was not "normal". We couldn’t run the HTTP flow.

So, we started with straightforward browser journeys in a single OCP pod. The first iterations looked fine. Then, about eight minutes into the test execution, browsers began to fall over. k6 would lose its DevTools connection and abort iterations mid-flow. The logs were blunt about it:

  • "unexpected EOF … connecting to browser DevTools URL"
  • "use of closed network connection"
  • "closing the browser: context deadline exceeded"
  • "process ended: trace/breakpoint trap (core dumped)"

Not the errors themselves, but their cadence is what caught our attention. Re-running the same test on the same pod made the crash happen faster. Minutes became under a minute, then only a few seconds. Running in a fresh pod resulted in a longer run. Hot pod? Quick collapse. That pattern was telling us the culprit was zombie processes: orphaned Chromium children piling up and poisoning the next launch.

The single-pod investigation

We started looking at fixing pod hygiene. Now that the blocker was revealed, the solution was straight-forward. So, we rebuilt the image with tini - a small init process for containers - as PID1, to stop zombie processes from hanging around. That one change made all the difference, unblocking a full 2500 iteration execution on a single pod, with no mid-flow collapses and no DevTools sockets dying. The first complete execution!

When did the blocker move from pod hygiene to resources?

As we nudged concurrency up by increasing the number of virtual users (VUs) per run, the failure pattern changed. With the process hygiene sorted, runs no longer died early.

  • At 25 VUs, the pod held steady with zero browser errors.
  • At 35 VUs on 12 CPUs, we saw just 50 browser errors, a manageable threshold.
  • Scaling to 40 VUs on the same pod crashed the execution.

Even after bumping the single pod to 12 CPUs, raising VUs much further became inefficient.

At that point the limiter moved from process hygiene to resources: CPU limits, not Chromium processes, was now our ceiling.

Once we found the single-pod breaking point, we stopped scaling vertically.

Why horizontal scaling beats vertical

Rather than keep pushing one pod harder, we scaled out. The k6 Operator enabled us to run the same test split across multiple pods, each with lower concurrency, so no single pod had to be a hero. Once we moved onto the k6 Operator and ran distributed tests, we didn’t hit the zombie process issue again. Every shard ran on a fresh pod with a clean lifecycle, so accumulation between retries never started.

What spread actually worked (from our run notes)

MetricSingle Pod (Vertical)Distributed (Horizontal)
Pods18 (Operator parallelism: 8)
CPUs per pod8 → 12 (tested)8
VUs per pod25–3525
Total sustained VUs35 (max stable)~200
Iterations completed2,500 successful journeys (after tini fix) in ~1h2,433 successful journeys in ~12 min
Errors per pod~50~7–8 (stable across pods)
RPS (requests per second)~1–2 rps~8 rps average, spikes >10 rps

After shifting to horizontal scaling with k6 distributed tests, stability improved. Each pod ran fewer browsers, but collectively returned higher throughput - roughly 2x the traffic we were seeing in production for this journey, even with browser-only k6 performance tests in a constrained pre-prod environment!

Keeping pods modest and scaling horizontally gave us our sweet spot of 25 VUs/pod × 8 pods × 8 CPUs/pod.

We knew this approach was costly. Browser-only performance testing aimed at reaching HTTP targets is naturally resource heavy, and we ultimately hit an internal cluster boundary - a practical ceiling on how much CPU and memory we could consume during our tests. In other words: you can’t scale forever.

The execution profile that gave us confidence (with code snippets)

A test bundle (archived k6 script) which landed in a ConfigMap, a TestRun to set parallelism (telling how many simultaneous pods will run), and each pod executing the journey with low concurrent browsers. We used ConfigMaps instead of a volume claim to serve tests to pods, as per the k6 Operator docs for executing scripts with the TestRun CRD, and we used the archive format because our tests are multi-file (script + helpers).

The image has k6 with browser support and tini when we run single-pod mode; on the k6 Operator, the clean pod lifecycle meant we didn’t need to change anything further.

Creating the k6 archive and ConfigMap for TestRun

Below are the commands we used to package the archive and wire it into the Operator.

cd <repo>/tests/k6
k6 archive k6-script.ts           # produce archive.tar with all deps
oc project <your-namespace>       # switch to your namespace
kubectl delete cm k6-script       # if configmap already exists, delete before re-creating
kubectl create cm k6-script --from-file=archive.tar   # create configmap from archive
jsx
Sanitised OCP TestRun skeleton

We set the following browser args in the runner: no-sandbox,disable-gpu,ignore-certificate-errors,no-zygote,renderer-process-limit=1.

These are stability levers for headless Chromium in containers. Disabling the zygote and limiting renderer processes reduce process churn and thread count, which helps avoid hitting per-pod task limits under load. In our case the dramatic improvement came from installing tini on the single-pod image; the args helped keep each browser instance predictable while we scaled.

apiVersion: k6.io/v1alpha1
kind: TestRun
metadata:
  name: k6-browser-test-<timestamp>     # e.g. k6-browser-test-2025-10-22-1500
  namespace: <your-namespace> 
spec:
  parallelism: 8                         # shards/pods (scale horizontally)
  separate: false
  quiet: 'false'                         # explicit: avoid default pause/silent surprises
  paused: 'false'
  script:
    configMap:
      name: k6-script              # created via `k6 archive k6-script.ts`
      file: archive.tar
  arguments: "--summary-export=/tmp/summary.json"
  runner:
    # Public example image with browser support;
    image: grafana/k6:master-with-browser
    env:
      - name: k6_BROWSER_DEBUG
        value: "false"
      - name: k6_BROWSER_ARGS
        value: "no-sandbox,disable-gpu,ignore-certificate-errors,no-zygote,renderer-process-limit=1"
      - name: BASE_URL
        valueFrom:
          secretKeyRef:
            name: <your-secret>
            key: BASE_URL
    resources:
      requests:
        cpu: "4"
        memory: "8Gi"
      limits:
        cpu: "8"
        memory: "16Gi"
    volumeMounts:
      - name: tmp
        mountPath: /tmp
    volumes:
      - name: tmp
        emptyDir: {}
jsx
k6 script options + thresholds hybrid example

We’ve defined our load models and custom thresholds as part of the script options object. This place provides automatic version control and allows for easy reuse, allowing you to modularise your script. Other places to define your options would be in CLI flags, environment variables or in a configuration file.

In the lifecycle of a k6 test, the code below runs as part of the init stage, which is required to prepare the test and initialise the test conditions. Code in the init context always executes first and runs once per VU. The below is a hybrid scenario example, where we define a small browser smoke test and a ramped HTTP load.

// k6 options & custom thresholds(sanitised, minimal)
// init context
const FE_PAGE = {
  settle_ms: new Trend('fe_route_settle_ms'),
};

const HTTP = {
  duration: new Trend('http_step_duration_ms'),
  ttfb: new Trend('http_step_ttfb_ms'),
  receive: new Trend('http_step_receive_ms'),
};

const THRESHOLDS = {
  browser_errors: ['count==0'],
  'fe_route_settle_ms{page:vop_flow}':          ['p(95)<2000'],
  'fe_route_settle_ms{page:pay_form}':          ['p(95)<2000'],
  'fe_route_settle_ms{page:confirm}':           ['p(95)<2000'],
  'fe_route_settle_ms{page:payment-processed}': ['p(95)<5000'],
  // HTTP examples
  'http_step_ttfb_ms{step:payees,method:GET}': ['p(95)<500'],
  'http_step_duration_ms{step:confirm,method:POST}': ['p(95)<1200'],
};

export const options = {
  thresholds: THRESHOLDS,
  scenarios: {
    browser_probe: {
      executor: 'shared-iterations',
      vus: 3,
      iterations: 3,
      startTime: '30s', // starts after the HTTP load is already in motion
      exec: 'browserProbe',
      gracefulStop: '5m',
      options: { browser: { type: 'chromium' } },
    },
    http_ramp10: {
	    executor: 'ramping-arrival-rate',
	    timeUnit: '1s',
	    preAllocatedVUs: 80,
	    maxVUs: 120,
	    startTime: '0s',
	    stages: [
	      { target: 3, duration: '10s' },
	      { target: 6, duration: '10s' },
	      { target: 10, duration: '40s' },
	    ],
	    gracefulStop: '10s',
	    exec: 'default',
	  }
};
jsx
Setup()

The setup() function is also running as part of the init context. This is where we do our backend login, when we run in higher environments, if an API_BASE_URL is available and return a sessionsByCard object. Both flows can then attach the correct session to the requests or to the browser context.

export function setup() {
  if (!RUN_HTTP) return {};
  const apiBaseUrl = __ENV.API_BASE_URL;
  if (!apiBaseUrl) return {};
  const sessionsByCard = buildSessionsByCard(apiBaseUrl); // sanitised helper
  return { sessionsByCard };
}
jsx
Browser journey shape (Playwright-style selectors + route settle)

The browser journey is part of the VU stage, which runs over and over through the test duration. This is where the similarity with Playwright code is striking, the only twist in our case of a single page application, is that we are wrapping route transitions in measureRouteSettle, which records how long each page takes to settle, and that feeds into the fe_route_settle_ms custom Trend defined above, in the init context.

// browserProbe(): core shape of the journey (sanitised)
// VU stage
export async function browserProbe(data: any) {
  const { scenario } = selectScenario();
  const sessionId = data?.sessionsByCard?.[scenario.card_number]?.sessionId ?? null;

  let context: any, page: any;
  try {
    context = await browser.newContext();
    await seedAspNet(context, BASE_URL, sessionId);
    page = await context.newPage();

    // 1) PAYEES
    const res = await page.goto(`${BASE_URL}${ROUTES.payees}`, { waitUntil: 'networkidle' });
    const perf = await collectPerf(page);
    if (perf.load != null) FE_PAGE.load.add(perf.load, { page: 'initial' });
    if (res && res.status() >= 400) throw new Error(`payees GET ${res.status()}`);

    const payFirst = page.getByRole('button', { name: 'Pay', exact: true }).first();
    await payFirst.waitFor({ state: 'visible', timeout: 60_000 });

    // 2) VOP SUCCESS
    await payFirst.click();
    await measureRouteSettle(page, 'vop_flow', () =>
      page.getByLabel('Payee verification: Success').waitFor({ state: 'visible', timeout: 60_000 }),
    );

    // 3) PAY FORM
    const next = page.getByRole('button', { name: 'Continue', exact: true });
    await next.click();
    await measureRouteSettle(page, 'pay_form', () =>
      page.getByRole('heading', { name: 'Make a payment', exact: false }).waitFor({ state: 'visible', timeout: 60_000 }),
    );

    // (inputs + confirm + processed)
    // … same pattern with measureRouteSettle('confirm') and ('payment-processed')
  } finally {
    try { await page?.close(); } catch {}
    try { await context?.close(); } catch {}
  }
}
jsx
HTTP flow

On the HTTP side, we are mirroring the browser journey as closely as possible, step by step. Each of the requests run in a group() block, and recordHttpPerf attaches the timings we care about: TTFB, total duration, download time, all tagged by step and method.

This code is also part of the VU stage in the lifecycle of a k6 test. The default() function gets executed by each VU from start to end in sequence. When the VU reaches the end of the function, the whole process restarts, looping back to the start and executing the code all over again.

// httpMakePayment(): minimal GET + POST examples (sanitised)
// VU stage
export default function httpMakePayment(data: any) {
  if (!RUN_HTTP) return;

  const { scenario } = selectScenario();
  const sessionId = data?.sessionsByCard?.[scenario.card_number]?.sessionId ?? null;
  const headers = feHeaders(sessionId);

  group('GET payees', () => {
    const r = http.get(`${BASE_URL}${ROUTES.payees}`, { headers, redirects: 10 });
    recordHttpPerf(r, 'payees', 'GET');
    check(r, { '200': x => x.status === 200 });
  });

  group('POST pay', () => {
    const form = new FormData();
    form.append('amount', scenario.amount);
    form.append('paymentDate', today());
    const h = mergeHeaders({ 'Content-Type': `multipart/form-data; boundary=${form.boundary}` }, sessionId);
    const r = http.post(`${BASE_URL}${ROUTES.pay}`, form.body(), { headers: h, redirects: 10 });
    recordHttpPerf(r, 'pay', 'POST');
    check(r, { '200/302': x => x.status === 200 || x.status === 302 });
  });
  // ..
}
jsx

The quiet win

The last test execution in pre-prod didn’t just “stay up.” It proved we could sustain high concurrent browser journeys across multiple pods with low error rates. This capacity signal was what we needed to release with confidence. And surprisingly, even with browser-only k6 tests, we reached twice our normal production volume on this journey.

What’s next?

As backend access is re-enabled in lower environments, we’ll fold these flows back into a hybrid setup with more journeys, broader coverage and cheaper scale.

Key takeaways

  • Scale sideways, not upwards, when vertical limits hit.
  • Browser-only load testing with k6 Operator can still validate UX under real load.
  • Small container hygiene fixes (like tini) can unlock big reliability gains.

This is our team’s success story with k6 load tests running under many constraints. We started with very low hopes that a browser-only approach could help validate backend targets, but we kept going and proved that teamwork, collaboration, and the right tooling can take you a long way, even under tight deadlines and pressure.

With special thanks to Marco Antonio Blanco (Staff DevSecOps Engineer), for his collaboration on the scaling strategy and Operator setup.

But wait - there's more.

Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.

Insights

Perspectives on AI in engineering, product development, and strategy, for enterprise executives.

Community

Deep dives and tutorials by engineers, for engineers.

Insight, imagination and expertly engineered solutions to accelerate and sustain progress.