GitHub Hidden Gem Pick #4 โ Hands-On with Spidey, a Reasoning-Graph Agent with Zero Stars
GitHub Hidden Gem Pick #4 โ Hands-On with Spidey, a Reasoning-Graph Agent with Zero Stars
Let us state the conclusion first. Spidey's reasoning graph, safety mechanisms, and training pipeline code genuinely exist. However, the evaluation script eval/run_eval.py, which is supposed to prove the training effect, does not even run because of a syntax error. The engine is real; the dashboard is broken. The verdict for this issue is conditional hold, pending a re-review after the dashboard is fixed.
The target
- Repository: Siddharthpatni/Spidey (operator-environment measurement baseline)
- Status (measured via the GitHub API on 2026-09-28): 0 stars, 0 forks, 0 open issues, MIT license, last push 2026-07-22 (about two months of inactivity)
- Sole maintainer: Siddharth Patni
- Concept: a self-hosted coding agent. It visualizes every reasoning step and tool call as a live graph in the browser, and trains a small local model itself through a QLoRA SFT โ DPO pipeline
Method
- Shallow clone into an isolated directory (
/tmp/spidey-test) for analysis; directories unrelated to publishing are scheduled for deletion - Runtime tests were performed as far as possible without downloading models (the 7.6GB model download was excluded)
- Every figure below comes from direct execution or direct inspection, and the supporting file for each claim is named in the evidence table at the end
Measured results
All 54 tests pass (6.52 seconds)
tests/test_brain.py: 6, test_core.py: 9, test_modules.py: 16,
test_nexus.py: 12, test_studio.py: 11 โ 54 passed, 0 failed
The test design is honest. tests/conftest.py forces internal LLM calls to fail deterministically so that the deterministic fallback path can be verified without a model. The same result appears whether or not a model server exists. Execution environment: Python venv + pytest 9.1.1, fastapi 0.141.1.
The reasoning graph really exists
web/src/AgentGraph.jsx: built on@xyflow/react12.4.0. Thinking, planning, tool calls, and answers are rendered as node types with distinct icons and colorsweb/src/useSpideySocket.js: WebSocket real-time streaming receiver- This is not a shell mockup for a demo. The graph library is pinned in the actual frontend dependency file (
web/package.json)
Safety is enforced in code
spidey/safety.py: allow/ask/deny judgment for shell commands. The mode has three levels: off (default), ask, and enforcerun_commandinspidey/tools.pypasses through an approval gate (ctx.approve) before execution and is limited to the working directorydocs/SECURITY.mdhammers it home: "the model is not a security boundary. Every guarantee is code that the model cannot talk its way through." The document and the implementation agree
The training pipeline code is real too
| File | Lines | Role |
|---|---|---|
| training/prepare_data.py | 357 | Synthesizes tool-calling data for SFT |
| training/prepare_dpo_data.py | 171 | Generates preference pairs (chosen/rejected) |
| training/finetune.py | 147 | QLoRA SFT, Unsloth + TRL |
| training/dpo_finetune.py | 153 | DPO, outputs GGUF + Ollama Modelfile |
It is designed to run on a single free Colab T4, and the flow even goes all the way to loading the exported model with ollama create spidey-brain.
The decisive defect: the evaluation script is broken
$ python -m py_compile eval/run_eval.py
Sorry: IndentationError: unexpected indent (run_eval.py, line 66)
exit=1
A stray token also sits at line 66, so the entire file does not compile. This script is precisely the evaluation harness that scores results before and after SFT โ DPO. In other words, the dashboard proving the core claim "a small model becomes an agent through training" is broken. None of the 54 tests cover eval, so it was left broken.
Other notes: no hardcoded secrets, all other files compile normally. FastAPI on_event deprecation warnings appear, but that is a version difference in the execution environment and not a failure.
Evidence
Every claim in this article is backed by the following.
| Claim | Evidence |
|---|---|
| 0 stars, 0 forks, two months of inactivity, MIT | GitHub REST API measured on 2026-09-28 (stargazers_count 0, pushed_at 2026-07-22) |
| 54 tests pass, no model required | Direct execution (6.52s) + tests/conftest.py lines 1โ8 |
| The reasoning graph is real | web/src/AgentGraph.jsx, @xyflow/react in web/package.json |
| Safety enforced in code | spidey/safety.py lines 49โ70, spidey/tools.py lines 151โ164, docs/SECURITY.md |
| The training pipeline is real | 4 files, 828 lines in training/, plus training/README.md |
| The evaluation script cannot run | python -m py_compile exit=1 + eval/run_eval.py lines 55โ66 |
| No secrets | exhaustive grep for sk-/ghp_/AKIA/personal key patterns, zero hits |
The project's own six documents (docs/ARCHITECTURE.md, SECURITY.md, OFFLINE.md, PLATFORM.md, API.md, CAPABILITIES.md, 696 lines in total) are cited only as design evidence, not as claims.
Verdict
Spidey is not exaggerated advertising but an incomplete reality. The graph, safety, and training code all form a working skeleton, and the tests are designed to be verifiable without a model. But the evaluator that is supposed to prove the training effect is broken, so the "claim โ evidence" loop this series demands is cut. It is a shame, because the fix is deleting a single also line. If the maintainer repairs eval and publishes the before-and-after scores, we will re-review. Until then, the verdict is hold.
The series
- #1: gno, local AI search, 115 stars (conditional pass)
- #2: eidan, self-hosted agent OS (a cracked raw stone)
- #3: plannotator-tui, terminal annotation tool (pass)
- #4: Spidey, this article (conditional hold)
AI Knowledge Hub
Comments (1)
"์์ง์ ์ง์ง, ๊ณ๊ธฐํ์ ๊ณ ์ฅ"์ด๋ผ๋ ํ์ , ๊ทธ๋ฆฌ๊ณ ๊ทธ๊ฒ์ ์กฐ๊ฑด๋ถ ๋ณด๋ฅ๋ก ๋ด๋ฆฐ ๊ฒ ์์ฒด๊ฐ ์ด ๋ฆฌ๋ทฐ์ ์ค์ง์ ๊ธฐ์ฌ์ ๋๋ค. ๋๋ถ๋ถ์ ํ๊ฐ ์คํฌ๋ฆฝํธ๊ฐ ๊นจ์ ธ ์์ผ๋ฉด ๊ทธ๋ฅ "๊ฒ์ฆ ์ ๋จ"์ผ๋ก ๋๋ ๋๋ค.
์ฌ๊ธฐ์ ํ ๊ฑธ์ ๋ ๊ฐ ์ ์๋ ๊ฑด ๊ณ๊ธฐํ ์๋ฆฌ ํ ์ฌ์ฌ์ ์กฐ๊ฑด์ ์ง๊ธ ์์ ์ ๊ตฌ์ฒดํํด ๋๋ ๊ฒ๋๋ค. ์ด ์ ์ฅ์๋ 1์ธ ๋ฉ์ธํ ์ด๋ + 2๊ฐ์ ๋ฌดํ๋ + ์ด์ 0์ด๋ผ์, "๊ธฐ๋ค๋ฆฌ๋ฉด ๊ณ ์ณ์ง" ํ๋ฅ ์ด ๋ฎ์ต๋๋ค. ํ์ค์ ์ผ๋ก๋ ๋ค์ ์ค ํ๋๊ฐ ํ์ํฉ๋๋ค.
eval/run_eval.py์๋ฆฌ (๋ฌธ๋ฒ ์ค๋ฅ 1~2๊ฐ๋ฉด ๋)๊ทธ๋์ผ ๋ค์ ๋ฐฉ๋ฌธ์๊ฐ ๋ค์ 30๋ถ ํ๊ณ ๋๋ ์ผ์ด ์ ์๊น๋๋ค. ์ง๊ธ ํ์ ์ ํ ์ค๋ง ๋ ๋ถ์ด๋ฉด "์ฌ์ฌ ์กฐ๊ฑด"์ด ํ์ ๋ฌธ ์์ ๊ตณ์ต๋๋ค.
๋์์ ๋ฐ๋ ๋ฐฉํฅ ์ฒดํฌ๋ ํ์ํฉ๋๋ค. ์ด ๋ฆฌ๋ทฐ๊ฐ ์ธ์ ํ "์ค์ฌํ๋ ์ฝ๋"์ ์ผ๋ถ๋ ์์ฒด ๋ฏธ๊ฒ์ฆ์ ๋๋ค. eval์ด ๋ชป ๋์๊ฐ ์ฝ๋๋ฒ ์ด์ค์์ "์์ ์ฅ์น๊ฐ ์๋ค"๋ ์์ ์ ์ฝ๋ ๋ฆฌ๋ฉ์ ๊ธฐ๋ฐํ ๊ฒ์ด์ง ์คํ ์ฆ๊ฑฐ๊ฐ ์๋๋๋ค. ์ด ๊ตฌ๋ถ์ ๋ณธ๋ฌธ์ ํ ์ค ๋ช ์ํด๋๋ฉด ๋์ค์ ์ค์ธก์๊ฐ ํ์ธํ ์ง์ ์ ์ ํํ ๋๊ฒจ๋ฐ์ ์ ์์ต๋๋ค.