Browser-based vector search: fast, private, and no backend requiredBrowser-basedvectorsearch:fast,private,andnobackendrequiredBrowser-basedvectorsearch:fast,private,andnobackendrequired
Ryan Roemer
10 Feb 2026
Share
Exploring the tools, tradeoffs, and real-world benefits of running vector search fully in the browser.
At Nearform, semantic similarity is a staple for many of our AI solutions. Typically, this means reaching for cloud-based datastores like pgvector (Postgres) or Elastic/OpenSearch. But as browsers gain AI-native features like WebGPU support and embedding models get smaller, the "backend-only" rule for vector search is changing.
We’ve been experimenting with moving the entire search stack (indexing, vector similarity, retrieval, and more) into the browser. The results are impressive — near-instantaneous in-memory queries with no network roundtrips.
Even beyond pure search performance, there are several advantages and features that browser-based vector search can potentially unlock.
Offline capability: After your initial data loads, the app can run entirely offline (perfect for progressive web apps).
Privacy: Searches never leave the browser, useful for compliance situations and regulated industries.
Simplicity: No database infrastructure to provision or scale.
While not a fit for massive datasets, a surprising amount can be done with plain old JavaScript in a browser. So what tools are best to pull this off? We evaluated several options and landed on Orama.
Choosing a client-side vector store: Orama
Several datastores support client-side vector search, such as RxDB, SQLite, and IndexedDB wrappers. After a review of these and other potential candidates, we settled on Orama JS.
We’ll note upfront that Orama and Nearform have a deep connection. Orama began as an open source project named Lyra, created by Michele Riva at Nearform in 2022. Michele would then go on to found Orama in 2023 based on the Lyra project, with support from Nearform.
History and goodwill aside, we picked Orama because of its powerful features and web-optimised capabilities, including the following:
It’s lightweight: Orama has no dependencies and is just 80K minified.
It’s fast: The initial database loads in our demo take about 70-100ms and vector query time is consistently in the 5-10ms range.
It’s flexible: Orama was originally launched with full-text (lexical) query support with advanced features like facets, geosearch, grouping and more. Orama can query via the full-text or vector features in addition to doing both in hybrid queries.
It’s developer-friendly: Clean APIs for creating databases, loading documents, and querying.
Although it’s out of scope for our article today, Orama also works really well in a traditional backend or cloud service.
Finding a web-sized embedding model
We chose the gte-small model via @xenova/transformers for our embedding model. It strikes a nice balance for size, speed, and relevance features — and most importantly, it works well in a browser:
Capability: It has 384 dimensions and supports up to 512 max tokens per input. While this is less dimensions and context window than e.g., OpenAI’s text-embedding-ada-002 (1536 dimensions with 8192 max tokens), it is sufficient to produce semantically meaningful text chunks.
Size & speed: It’s a small, quick download at ~30mb and creating embedding results is fast (20-30ms in our demo case) and can compute via CPU or WebGPU.
Hooking this up to our posts content is straightforward:
import{ pipeline }from"@xenova/transformers";const extractor =awaitpipeline("feature-extraction","Xenova/gte-small");const text ="this is some text!";const output =awaitextractor(text,{pooling:"mean",normalize:true});console.log(Array.from(output.data));// =>[-0.04172841086983681,-0.0032770344987511635,0.04789912328124046,...381 more items
]
jsx
From raw text to embeddings
Our data
With our embeddings model in hand, we’ll find a good collection of text data to show off what we can do in a browser. For this post, we chose Nearform’s blog posts and case studies, which we aggregated in an out of band process into a posts.json endpoint. The data shape includes the following (with some omissions):
The field that we’re going to focus on — and ultimately perform our vector search on — is content, which is an array of strings that comprise sections of a Nearform blog post or case study.
There are about 900 articles in the data set. The average number of words in the content field of an article is around 1,200 words, with our largest article maxing out at 7,600 words. All in all, it’s not that big of a data set — 7.8 mb uncompressed and unminified — but it’s also not trivially small. Those curious about the raw data can find it here: posts.json.
Splitting and vectorizing chunks
So how do we vectorise this content? Using the gte-small tokeniser, our posts have on average around 1,900 tokens with a max content of 13,000 tokens. However, gte-small can only produce embeddings for up to 512 tokens, so we’ll need to break the content down into smaller segments.
Generally speaking (beyond our hard constraint of 512 tokens) there are many strategies for optimal text chunking. To keep things simple, we’ll implement two chunking sizing strategies for this demo: 512 tokens and 256 tokens. We won’t dive deeper here into the details of creating the embeddings, but we’ll leave you with two salient points:
We use Nearform’s own llm-splitter library for chunking. It’s fast, works in the browser, and you can read more about it in our introductory article.
The embedding creation script is here: embeddings.js.
Putting this all together, we can split up a post content field into chunks with embeddings in files that resemble something like:
{"<SLUG_01>":{"chunks":[{"start":0,"end":2132,// text positions in `content`"embeddings":[-0.04,-0.0035,0.047,...]},{"start":2095,"end":4240,"embeddings":[-0.0110,0.141,-0.144,...]},]},"<SLUG_02>":{...},...}
jsx
Optimising the embeddings
The astute reader will note the above JSON format isn’t actually how our embeddings JSON is structured. We make a quick transformation to significantly slim down the files, and size is an important performance consideration because our in-browser vector search must download embeddings on initial load.
We quantise the embeddings field, converting long floats of [FLOAT, FLOAT, ...] shape into much smaller integers that we bound with a minimum and maximum float range. Our real embeddings field looks like { values: [INT, INT, ...], min: FLOAT, max: FLOAT }, produced by a quantizeEmbedding function (that can be reversed with dequantizeEmbedding). This optimisation makes using the data slightly more complicated, but it reduces our JSON file size by 75%, which is much easier to work with for our demo. The optimisation has a negligible impact on the embedding float precision and resulting query results — we examined the 4.5 million floats for this data set and found the average transformation had less than a 0.08% value delta for any given embedding number.
In the real world, there are much more efficient and sophisticated quantisations you can do, as well as much better transport data formats than the prettified JSON we use to keep things simple. Here are the final embeddings files we’ll put to use in our demo: posts-embeddings-512.json and posts-embeddings-256.json. (Like posts.json, these files are pretty large at 2mb and 4mb respectively).
Putting it all together — schema, loading, querying
We have all the pieces and technical decisions in place — let’s go build a vector search application. The main “search” part of this is an Orama database of our chunks. We can then find the most semantically relevant chunks to a given query and return information about the chunk and the article it belongs to.
A schema for our chunks
Creating an Orama database is straightforward, but there are a few limitations in schema definitions. Conceptually, we have a post that has an array of chunks with text identifiers and embeddings. Unfortunately, while Orama supports many different types, it does not support nested arrays of custom objects like our chunks.
Thus, we choose a flattened chunks database instead of a nested posts database — and then denormalize and repeat some information from the posts data to be able to filter and query. This ends up looking like:
import{ create }from"@orama/orama";const db =create({schema:{// Post metadata for filtering.slug:"string",date:"number",postType:"string",categories:{primary:"string",},// Chunk data.start:"number",end:"number",embeddings:"vector[384]",// gte-small has 384 dimensions},});
jsx
Loading our database
Next we load our posts and embeddings, combine them into relevant chunk records, and insert them all in our database. In addition to our dequantizeEmbedding optimisation wrapper, we also convert dates from "YYYY-MM-DD" format to Unix epoch date numbers for querying. Let’s put that all together:
constURL_BASE="<https://raw.githubusercontent.com/nearform/joyce/main/public/data>";constget=(url)=>fetch(`${URL_BASE}/${url}`).then((res)=> res.json());const posts =awaitget("posts.json");// We're just using the 512 token example. For our full demo, we use both.const chunkEmbeddings =awaitget("posts-embeddings-512.json");// Flatten and combine chunks.const chunks =Object.entries(chunkEmbeddings).flatMap(([slug,{ chunks }])=>{const post = posts[slug];return chunks.map((chunk)=>({// The post data for filtering and matching. slug,date:dateToNumber(post?.date),postType: post?.postType,categories: post?.categories,// The chunk data. With embeddings converted back to floats....chunk,embeddings:dequantizeEmbedding(chunk.embeddings),}));});// Insert all chunks into the Orama database.insertMultiple(db, chunks);
jsx
After the call to insertMultiple, we have an in-memory database that we can query.
Querying the database
We generate embeddings for a query text value that are placed in the Orama vector field. Our search give us score and rank record results according to semantic similarity:
import{ search }from"@orama/orama";import{ pipeline }from"@xenova/transformers";const extractor =awaitpipeline("feature-extraction","Xenova/gte-small");const opts ={pooling:"mean",normalize:true};// Create embeddings for our query.const query ="React testing";const queryEmbedding =awaitextractor(query, opts).then(({ data })=>Array.from(data),);// Get 10 results with the highest similarity score.const results =search(db,{mode:"vector",vector:{value: queryEmbedding,property:"embeddings"},limit:10,});// Raw hits: stored documents and scoreconsole.log(results.hits);
jsx
We’re doing semantic similarity search! The above code logs out raw JSON results of the top 5 chunk records by similarity score to the query embedding generated for “React testing.” For ease of viewing, here’s a prettified version:
If you were to review the text, the chunks found here would all be semantically relevant. We also see that our most common primary category is indeed “test.” Taking this to the next level, we can also do filter on any field our records. Let’s now do the same search, but limit to chunks from documents (1) with a primary category of “mobile” and (2) published from 2023-on. We do this with a where clause:
And with this work, we now have a respectable, flexible vector search!
Try it out: the live demo
The above steps are pretty much all you need for a complete vector search solution that runs fully in a browser! Try the demo at https://nearform.github.io/vector-search-web/. You can filter by type, category, date, and chunk size. And, you can dig into all the underlying data details by clicking the { } icon in the results header.
The drawback of on-device vector search is that you must load the data (at least initially) and fit it in a client-side database. Once you do, querying can be blazingly fast because you have no further network involved. Our demo has some timings logging (with the ?timings=true flag). Using the flag, we get some rough timings from our demo app:
Download embeddings 512: 300ms
Download embeddings 256 (a bigger file): 700ms
Download posts: 800ms
Create Orama database for 256-sized chunks: 200ms
Create Orama database for 512-sized chunks: 175ms
Most of these things can be parallelised, so we're looking at about 1-2 seconds of load time until the app is ready to be queried. After that, it’s really fast:
Converting a query into an embedding takes about 20-30ms
Searching Orama takes only about 6-7ms
With search times like these, the results feel virtually instantaneous in the UI!
Where can we go from here?
At this point, we’ve walked through a powerful, flexible, and fast vector search app that runs entirely in the browser. But this is just a starting point that opens up many other interesting possibilities.
Hybrid search
Orama is a fully capable lexical search tool with many additional features — many of which can be combined with vector search. For starters, here’s a simple example that combines BM25 keyword relevance with semantic vector similarity, with a custom scoring balance.
const results =search(db,{mode:"hybrid",term:"Next.js",// Full-text search termvector:{value: queryEmbedding,property:"embeddings"},limit:10,where:{date:{gte:dateToNumber("2023-01-01")},},// Adjust the balance between text and vector scoringhybridWeights:{text:0.7,// weight on BM25 keyword matchingvector:0.3,// weight on semantic similarity},});
jsx
The ability to combine these different search features gives us powerful relevance options for everything from finding better results for users to powering better data context upstream to applications.
On-device AI applications
We’ve already fit an embedding model in the browser, and it turns out that generative models can fit as well. While Full-scale LLMs are still too large for most browsers, we can turn to the neat world of Small Language Models (”SLMs”) and tools. Projects like web-llm can run open weight SLMs in a browser with hardware acceleration. And, the Chrome browser continues to integrate advanced AI functionality, like the Prompt API that enables inference with a local copy of the Gemini Nano model, and with quite impressive outputs for such a small model.
It’s a natural step to pair our vector search journey here with a SLM to be able to build things like a fully browser-based retrieval-augmented generation (”RAG”) application or a vision capable application. The possibilities are endless with the combination of device inputs (voice, text, photos, videos) and use cases like offline functionality, fully private on-device data, and more. We look forward to exploring some of these ideas in future articles, so stay tuned!
The future is big (and small)
Nearform has been at the forefront of modern web development for well over a decade. And we’ve been invested in the rise of AI as well, both in bringing new functionality and power to web and mobile applications as well as overhauling software engineering itself with the advent of AI native engineering. The emerging space of vector search and AI within the web browser is such a natural and exciting fit for us and the technology community.
Keep an eye on our posts for more discussion of AI applications, both big and small. And if you could use a trusted expert across all these worlds — web, mobile, vector search, language models, and more — get in touch and see if Nearform can help you ship something amazing.
But wait - there's more.
Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.
Insights
Perspectives on AI in engineering, product development, and strategy, for enterprise executives.