Neural Workspaces: A Knowledge Graph Built From Capitalized Words

Neural Workspaces is a repository of mine under github.com/yethikrishna. Its README calls it an AI-powered workspace capture and knowledge graph SaaS. The graph logic sits in backend/src/graph-engine/KnowledgeGraph.ts, and I read that file. This post covers that file only.

How a capture becomes a graph

buildGraphFromCapture takes a saved capture and extracts entities from its content. The extractor splits the text on whitespace, keeps words that match a single capital letter followed by lowercase letters, removes duplicates and takes the first five. Each becomes a CONCEPT node. The first 50 characters of the text always become one DOCUMENT node. A comment in the file says a production version should use an NLP library such as spaCy or NLTK.

Nodes are upserted on workspace, type and label, so the same word in two captures maps to one node. Then every pair of nodes gets a RELATED_TO edge. A new edge starts at weight 1.0, and seeing the same pair again adds 0.1 to the weight. That is a sensible way to let repeated co-occurrence strengthen a link.

What it gets wrong

A regex for capitalized words finds sentence starters as readily as names. It misses anything lowercase, any acronym, any phrase of two words and anything in a language without capitals. At most five concepts come out of a capture, so the graph is closer to a list of capitalized words than a map of ideas. Linking every pair also means each capture of six nodes creates 15 edges, and the edges are written one await at a time in a loop.

The return value is a smaller problem that I would fix first. The function reports edges as Math.pow(nodes.length, 2) / 2. For six nodes that prints 18, while the loop makes 15, because the count of pairs is n times n minus 1, over 2. It is an estimate dressed as a count.

The analytics and the export

getGraphAnalytics computes density as edges over the maximum possible edges, and average connections as twice the edges over the nodes. Both are the standard formulas. It then reports cognitiveLoad as the smaller of 100 and density times 50 plus average connections times 10. I could not find where those weights come from in the file, so that number is a design choice of mine, not a measured quantity, and the UI should not present it as one.

exportGraph has a type that allows json, csv and graphml. Only json is implemented. The other two fall through and return the raw node and edge rows, with a comment that other formats follow.

What I will do with it

My opinion: the structure around the graph is fine and the extractor is the weak part. Replace the regex with a real entity extractor, return the true edge count, and either implement csv and graphml or remove them from the type. The graph engine is a prototype, and I would describe it that way.

← back to the journal