Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

179 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Graphos

Context graph builder — uses the Language Server Protocol to extract code as a graph, consolidate knowledge into a context graph, and save it with your project — so you use fewer tokens per LLM call.

What Graphos Does

Graphos takes any folder of code, docs, papers, and images and builds a navigable knowledge graph with community detection. It produces interactive HTML, queryable JSON, and a plain-language audit report.

The key innovation: LSP to generate context graphs and create context graphs to optimise LLM context. Graphos connects to any language server. If a language has an LSP server, Graphos can extract its structure. That means:

  • TypeScript, JavaScript, Python, Go, Rust, Java, C#, Haskell, Erlang, Zig — all supported
  • Elm, PureScript, Idris, Agda — all supported
  • Your custom DSL with an LSP — supported
  • Every new language server that ships — automatically supported

Supported File Types

TypeExtensionsExtraction
Code.py.ts.js.jsx.tsx.go.rs.java.c.cpp.h.hpp.rb.cs.kt.kts.scala.php.swift.lua.zig.ps1.ex.exs.m.mm.jl.vue.svelte.dart.hs.lhsAST via tree-sitter + call-graph (cross-file for all languages) + docstring/comment rationale + LSP
Docs.md.txt.rst.adoc.orgConcepts + relationships + design rationale via LLM
Office.docx.xlsxConverted to markdown then extracted via LLM
Papers.pdfCitation mining + concept extraction
Images.png.jpg.jpeg.webp.gifLLM vision — screenshots, diagrams, any language
Video/Audio.mp4.mov.mkv.webm.avi.m4v.mp3.wav.m4a.oggTranscribed locally with faster-whisper, transcript fed into LLM extraction

Pipeline

detect() → extract() → build() → cluster() → infer() → analyze() → export()

Each stage is a pure function. No shared state, no side effects outside graphos-out/.

Architecture

src/Graphos/
├── Domain/ -- Pure types, no IO
│ ├── Types.hs -- Node, Edge, Extraction, Confidence
│ ├── Graph.hs -- Graph operations (add, merge, query, shortest path)
│ ├── Community.hs -- Leiden community detection
│ ├── Analysis.hs -- God nodes, surprising connections, suggested questions
│ └── Extraction.hs -- Extraction schema, validation
│
├── UseCase/ -- Orchestration, still pure
│ ├── Pipeline.hs -- Full pipeline orchestration
│ ├── Detect.hs -- File detection
│ ├── Extract.hs -- LSP extraction + Haskell stub fallback
│ ├── Build.hs -- Graph construction from extractions
│ ├── Cluster.hs -- Community detection
│ ├── Analyze.hs -- Analysis orchestration
│ ├── Report.hs -- Report generation
│ ├── Export.hs -- Export orchestration
│ ├── Query.hs -- Graph querying (BFS, DFS, shortest path)
│ └── Infer.hs -- Edge inference (community bridges, transitive deps)
│
└── Infrastructure/ -- IO boundary, all side effects here
├── LSP/
│ ├── Client.hs -- Connect to language servers
│ ├── Protocol.hs -- LSP JSON-RPC protocol types
│ └── Capabilities.hs -- Language server capability detection
├── FileSystem/
│ └── Watcher.hs -- File watching for --update
├── Export/
│ ├── JSON.hs -- graph.json output
│ ├── HTML.hs -- graph.html (interactive vis.js)
│ ├── Obsidian.hs -- Obsidian vault
│ ├── Neo4j.hs -- Cypher generation
│ ├── GraphML.hs -- GraphML for Gephi/yEd
│ ├── SVG.hs -- Static SVG export
│ └── Report.hs -- GRAPH_REPORT.md
└── Server/
└── MCP.hs -- MCP stdio server

Clean Architecture Principles

  1. Dependencies point inward: Domain ← UseCase ← Infrastructure. Domain knows nothing about LSP, IO, or any library.
  2. All domain logic is pure: Graph operations, community detection, analysis — all pure functions. Testable without mocks.
  3. LSP is an adapter: The domain doesn't know about LSP. It just receives extraction results. The LSP client adapter produces those results.
  4. Standard output format: graph.json for interoperability with visualization tools and queries.

Why LSP Instead of tree-sitter?

Aspecttree-sitterLSP (Graphos)
Language support25 hardcoded grammarsAny language with an LSP server
New languageAdd grammar + recompileJust install the LSP server
Semantic infoSyntax only (AST)Symbols, references, call hierarchy, type info
Cross-file refsSecond-pass inferenceNative via LSP references/callHierarchy
Hover/docsNot availableAvailable via LSP hover
MaintenanceGrammar per languageZero — LSP servers maintained by language teams
OfflineWorks without language serverRequires LSP server installed

Install

cabal install graphos

Or with stack:

stack install graphos

Language Server Requirements

Graphos auto-detects installed language servers. Install the ones you need:

# Common language servers (examples)
npm install -g typescript-language-server typescript # TypeScript/JS
npm install -g vscode-langservers-extracted # HTML/CSS/JSON
pip install python-lsp-server # Python
go install golang.org/x/tools/gopls@latest # Go
rustup component add rust-analyzer # Rust
cabal install haskell-language-server # Haskell

Ignore Patterns

Graphos honours .gitignore and .graphosignore files to exclude build artifacts, dependencies, and other irrelevant files. Use --ignore GLOB to add additional patterns at runtime.

.graphosignore

Create a .graphosignore file in your project root to declare patterns that should always be excluded. Syntax matches .gitignore:

# Exclude build outputsdist/
build/
target/
# Exclude large binary assets*.pdf*.mp4

Where it is read:.graphosignore is read from the scan root directory (the directory passed to graphos scan <DIR> or the directory argument to graphos), not the current working directory.

Match semantics:

  • Patterns match against scan-root-relative paths using normalized forward slashes
  • Backslashes are converted to forward slashes; . and .. path components are resolved
  • * matches within a single path component; ** matches across components
  • Leading / anchors to the scan root; leading **/ matches any prefix
  • Trailing / matches directories only; ! negates a pattern
  • Comments (#) and blank lines are ignored

--ignore flag

Pass additional patterns via the CLI. Can be repeated for multiple patterns:

graphos . --ignore "**/vendor/**" --ignore "*.log"

Patterns are merged with .gitignore and .graphosignore patterns. File-level patterns (e.g., *.log) match against the basename; directory patterns (e.g., vendor/) match against path components. CLI patterns are applied in addition to any patterns from .graphosignore files.

Extraction Fidelity Harness

The harness validates the fidelity of extraction against ground truth from the source files and gives users a path/taxonomy-driven subgraph facility. All three components are part of the standard graphos build and cabal test — no external interpreter or runtime is required.

ComponentPurposeInvocation
ImportEdgesSpecOn-disk oracle for imports edges (precision/recall, gap listings)cabal test --match ImportEdges
GraphCoverageSpecFile coverage accounting grouped by ignore-rule classcabal test --match GraphCoverage
graphos subgraphExtract a pattern-selected subgraph from a graph.jsongraphos subgraph --graph <g.json> --config <cfg.json> --out <out.json>

Exit codes: the Hspec specs pass with exit code 0 and fail with a non-zero exit code when precision/recall (imports) or any unexplained file (coverage) drops below the gate. The graphos subgraph command exits 0 on success, 1 when --config is required but missing or the config/graph files cannot be parsed.

ImportEdgesSpec

Scans a repository on disk, resolves every import/re-export specifier to a file, and compares the resulting pair set with the imports edges in a graph.json. It reports the ground-truth pair count, the graph edge count, and the precision/recall gaps as explicit MISSING/EXTRA pair listings. The spec fails when precision or recall drops below the threshold (default 0.99).

cabal test --match ImportEdges

GraphCoverageSpec

Compares the source files on disk with the files present in a graph.json and groups any missing files by the ignore-rule class that most plausibly explains them: root-anchored build output, depth-independent tooling, .gitignore, or unexplained. The spec fails when any file is unexplained, so the "unexplained" bucket can be fed back into gitignore parsing.

cabal test --match GraphCoverage

graphos subgraph

Extracts a subgraph from an existing graph.json by selecting core files from path patterns grouped into named subsystems, expanding a boundary tier of files that import a core file or are imported by one, and an external tier of package dependencies. Output conforms to the graph.json contract and is directly consumable via --graph (query/explain/neighbors). Every node carries tier/subsystem/layer metadata and every edge carries a provenance marker (source or derived).

Flags:

FlagDefaultDescription
--graph PATHgraphos-out/graph.jsonSource graph to extract from
--config PATH— (required)Subsystem patterns JSON
--out, -o PATHgraphos-out/subgraph.jsonOutput graph path
--boundary-hops N1Import-graph BFS depth for the boundary tier
--no-derivederive enabledDisable deriving imports edges from Import nodes
graphos subgraph --graph graphos-out/graph.json --config subgraph-config.json \
--out graphos-out/subgraph.json
graphos query "auth" --graph graphos-out/subgraph.json
graphos explain "RequestHandler" --graph graphos-out/subgraph.json
graphos neighbors "RequestHandler" --graph graphos-out/subgraph.json

Config schema (--config):

{
"subsystems": [
{ "name": "detect", "patterns": ["src/UseCase/Detect/**"] },
{ "name": "ignore", "patterns": ["src/Infrastructure/FileSystem/Ignore*"] }
],
"max_hops": 1,
"include_derived": true
}
# Full pipeline on current directory
graphos .# Specific folder
graphos ./my-project
# Directed graph (preserves edge direction)
graphos ./my-project --directed
# Skip visualization
graphos ./my-project --no-viz
# Incremental update (only changed files)
graphos ./my-project --update
# Watch mode
graphos ./my-project --watch
# Additional ignore patterns
graphos ./my-project --ignore "**/vendor/**" --ignore "*.log"# Query the knowledge graph (natural language)
graphos query "how does authentication work?"
graphos query "how does authentication work?" --dfs
graphos query "how does authentication work?" --budget 5000
graphos query "how does authentication work?" --graph path/to/graph.json
# Find shortest path between two nodes
graphos path "AuthModule""Database"
graphos path "AuthModule""Database" --graph path/to/graph.json
# Explain a node (show all connections)
graphos explain "RequestHandler"
graphos explain "RequestHandler" --graph path/to/graph.json
# List available LSP servers
graphos lservers
# Serve HTML visualization over HTTP
graphos serve --dir graphos-out --port 8080
# MCP server
graphos --mcp graphos-out/graph.json
# Export formats
graphos ./my-project --obsidian
graphos ./my-project --neo4j
graphos ./my-project --graphml
graphos ./my-project --svg

Query Options

FlagDefaultDescription
--dfsbfsUse DFS traversal instead of BFS
--budget N2000Token budget for query results
--graph PATHgraphos-out/graph.jsonPath to graph.json file

What You Get

graphos-out/
├── graph.html # Interactive graph - click nodes, search, filter by community
├── GRAPH_REPORT.md # God nodes, surprising connections, suggested questions
├── graph.json # Persistent graph - query weeks later without re-reading
└── cache/ # SHA256 cache - re-runs only process changed files

License

MIT

About

--- Graphos Knowledge Graph for LLMs Graphos builds a structured, queryable knowledge graph from any codebase, documentation, papers, and images — designed as persistent context that LLMs can traverse on-demand instead of loading entire repositories into a prompt window.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages