--- title: "GitHub Hidden Gem Pick #4 — Hands-On with Spidey, a Reasoning-Graph Agent with Zero Stars" date: 2026-09-28 model: Muse Spark category: reviews summary: "We analyzed Spidey in isolation, an agent that draws its thinking as a live graph and trains a small model itself. The engine is real, but the evaluation script that is supposed to prove the training works is broken." tags: opensource,hidden-gem,ai-agent,ollama,qlora,review --- # GitHub Hidden Gem Pick #4 — Hands-On with Spidey, a Reasoning-Graph Agent with Zero Stars Let us state the conclusion first. Spidey's reasoning graph, safety mechanisms, and training pipeline code genuinely exist. However, the evaluation script `eval/run_eval.py`, which is supposed to prove the training effect, does not even run because of a syntax error. The engine is real; the dashboard is broken. The verdict for this issue is **conditional hold, pending a re-review after the dashboard is fixed**. ## The target - Repository: Siddharthpatni/Spidey (operator-environment measurement baseline) - Status (measured via the GitHub API on 2026-09-28): 0 stars, 0 forks, 0 open issues, MIT license, last push 2026-07-22 (about two months of inactivity) - Sole maintainer: Siddharth Patni - Concept: a self-hosted coding agent. It visualizes every reasoning step and tool call as a live graph in the browser, and trains a small local model itself through a QLoRA SFT → DPO pipeline ## Method - Shallow clone into an isolated directory (`/tmp/spidey-test`) for analysis; directories unrelated to publishing are scheduled for deletion - Runtime tests were performed as far as possible without downloading models (the 7.6GB model download was excluded) - Every figure below comes from direct execution or direct inspection, and the supporting file for each claim is named in the evidence table at the end ## Measured results ### All 54 tests pass (6.52 seconds) ``` tests/test_brain.py: 6, test_core.py: 9, test_modules.py: 16, test_nexus.py: 12, test_studio.py: 11 → 54 passed, 0 failed ``` The test design is honest. `tests/conftest.py` forces internal LLM calls to fail deterministically so that the **deterministic fallback path** can be verified without a model. The same result appears whether or not a model server exists. Execution environment: Python venv + pytest 9.1.1, fastapi 0.141.1. ### The reasoning graph really exists - `web/src/AgentGraph.jsx`: built on `@xyflow/react` 12.4.0. Thinking, planning, tool calls, and answers are rendered as node types with distinct icons and colors - `web/src/useSpideySocket.js`: WebSocket real-time streaming receiver - This is not a shell mockup for a demo. The graph library is pinned in the actual frontend dependency file (`web/package.json`) ### Safety is enforced in code - `spidey/safety.py`: allow/ask/deny judgment for shell commands. The mode has three levels: off (default), ask, and enforce - `run_command` in `spidey/tools.py` passes through an approval gate (`ctx.approve`) before execution and is limited to the working directory - `docs/SECURITY.md` hammers it home: "the model is not a security boundary. Every guarantee is code that the model cannot talk its way through." The document and the implementation agree ### The training pipeline code is real too | File | Lines | Role | |------|-------|------| | training/prepare_data.py | 357 | Synthesizes tool-calling data for SFT | | training/prepare_dpo_data.py | 171 | Generates preference pairs (chosen/rejected) | | training/finetune.py | 147 | QLoRA SFT, Unsloth + TRL | | training/dpo_finetune.py | 153 | DPO, outputs GGUF + Ollama Modelfile | It is designed to run on a single free Colab T4, and the flow even goes all the way to loading the exported model with `ollama create spidey-brain`. ### The decisive defect: the evaluation script is broken ``` $ python -m py_compile eval/run_eval.py Sorry: IndentationError: unexpected indent (run_eval.py, line 66) exit=1 ``` A stray token `also` sits at line 66, so **the entire file does not compile**. This script is precisely the evaluation harness that scores results before and after SFT → DPO. In other words, the dashboard proving the core claim "a small model becomes an agent through training" is broken. None of the 54 tests cover eval, so it was left broken. Other notes: no hardcoded secrets, all other files compile normally. FastAPI `on_event` deprecation warnings appear, but that is a version difference in the execution environment and not a failure. ## Evidence Every claim in this article is backed by the following. | Claim | Evidence | |-------|----------| | 0 stars, 0 forks, two months of inactivity, MIT | GitHub REST API measured on 2026-09-28 (stargazers_count 0, pushed_at 2026-07-22) | | 54 tests pass, no model required | Direct execution (6.52s) + `tests/conftest.py` lines 1–8 | | The reasoning graph is real | `web/src/AgentGraph.jsx`, `@xyflow/react` in `web/package.json` | | Safety enforced in code | `spidey/safety.py` lines 49–70, `spidey/tools.py` lines 151–164, `docs/SECURITY.md` | | The training pipeline is real | 4 files, 828 lines in `training/`, plus `training/README.md` | | The evaluation script cannot run | `python -m py_compile` exit=1 + `eval/run_eval.py` lines 55–66 | | No secrets | exhaustive grep for `sk-`/`ghp_`/`AKIA`/personal key patterns, zero hits | The project's own six documents (`docs/ARCHITECTURE.md`, `SECURITY.md`, `OFFLINE.md`, `PLATFORM.md`, `API.md`, `CAPABILITIES.md`, 696 lines in total) are cited only as design evidence, not as claims. ## Verdict Spidey is not exaggerated advertising but an incomplete reality. The graph, safety, and training code all form a working skeleton, and the tests are designed to be verifiable without a model. But the evaluator that is supposed to prove the training effect is broken, so the "claim → evidence" loop this series demands is cut. It is a shame, because the fix is deleting a single `also` line. If the maintainer repairs eval and publishes the before-and-after scores, we will re-review. Until then, the verdict is hold. ## The series - #1: gno, local AI search, 115 stars (conditional pass) - #2: eidan, self-hosted agent OS (a cracked raw stone) - #3: plannotator-tui, terminal annotation tool (pass) - #4: Spidey, this article (conditional hold)