ContextWeave AI: A RAG Core, and a WebGPU Panel That Is a Timer

ContextWeave AI is a repository of mine under github.com/yethikrishna. Its README lists a context weaving engine, ChromaDB vector search, multi-LLM support and WebGPU acceleration. I opened backend/src/services/contextWeaver.ts and frontend/src/components/WebGPUAccelerator.svelte to see which of those the code backs. This is my reading of those two files only, not of the whole repo.

What the weaving engine does

embedDocument loads a document, splits its text into chunks of 1000 characters with 200 characters of overlap, gets an embedding for each chunk and adds them to a ChromaDB collection named after the context. Embeddings come from OpenAI's text-embedding-ada-002, one call per chunk, all fired at once with Promise.all.

weaveContext embeds the query, asks Chroma for the 5 nearest chunks by default, pastes them into a prompt and asks the chat model for an answer, telling it to say so if the context does not contain the answer. The default model is gpt-4. That is a plain retrieval-augmented generation loop, and it is a sound shape for a first version.

What is thin

The chunker cuts on character counts, so it can split a sentence or a word in half. The Promise.all over every chunk has no concurrency limit, so a long document means a burst of embedding requests at once. And temperature is read as options.temperature || 0.7, which means a caller who passes 0 for deterministic output silently gets 0.7.

In this file the only model client is OpenAI. The README says OpenAI and Anthropic. Anthropic support may live elsewhere in the repo, and I did not check, so I will not claim it either way.

The WebGPU part

The service method webgpuAcceleratedEmbedding returns the object {accelerated: true, device: 'webgpu'} and nothing else. Its own comment says it would interface with client-side WebGPU. The Svelte panel does detect whether the browser has a GPU adapter and shows its vendor and architecture, which is real. But the Run Accelerated Embedding button awaits a 1.5 second timer, with a comment saying it simulates the operation. The line telling the user that inference is accelerated 8 to 12 times is fixed text in the template, not a measurement.

My opinion: this is a mockup of a feature, and the README should say so. A panel that reports a speedup nobody measured is worse than no panel, because someone will believe it. I would either build the WebGPU path with a benchmark behind the number, or relabel the panel as a prototype and move the claim out of the feature list.

What I will do with it

The retrieval core is worth keeping and is the part I would harden first: sentence-aware chunking, a concurrency limit on embeddings, and a temperature check that treats 0 as a value. The WebGPU claim comes out of the README until there is code behind it.

← back to the journal