--- title: "How Far Can a 7.5B Model Go as a Personal Assistant? A Real-World Test of 17 Tasks with SuperGemma" date: 2026-10-02 model: supergemma category: reviews summary: I put a 7.5B model running on 8GB of VRAM through real personal-assistant tasks. Document organization, summarization, and shell commands were accurate, but there were clear walls in its ability to choose tools on its own (60%) and to hold a multi-step task to the end. tags: "local AI, personal assistant, agent, SuperGemma, 8GB VRAM, Ollama, real-world test" author_type: human --- To get to the conclusion first: **a 7.5B-class local model is enough as a "personal assistant that answers when you ask," but it still falls short as an "assistant that judges for itself and carries a multi-step task all the way through."** Document organization, summarization, sentence writing, and shell-command knowledge passed with essentially no errors. But as soon as tool calls were attached and tasks ran two or more steps, a clear point of collapse emerged. This post is a measured record of exactly where that boundary lies and which settings push it back. --- ## Test Environment (measured in the operator's environment) All the numbers below are **measured values from a single environment.** They will differ with different hardware and settings. | Item | Value | |---|---| | GPU VRAM | 8,192 MiB | | RAM | 31 GiB | | Ollama | 0.33.3 | | Model | `supergemma-e4b-q4km-63k` | | Architecture | gemma4 (effective 4.5B / raw 7.5B multimodal) | | Quantization | Q4_K_M, 5.3 GB | | Context ceiling | 131,072 | | thinking feature | available, but **fully disabled** for this entire test | The model parameters were fixed as follows. ``` PARAMETER num_ctx 65500 PARAMETER num_predict 3072 PARAMETER temperature 0.4 PARAMETER top_p 0.9 PARAMETER top_k 20 ``` Speed measured at **about 75–83 tok/s**, and loading a 64k context occupied **4,815 MiB** of VRAM. It runs 100% on GPU within 8GB. --- ## Area 1. Pure Quality — this part almost passes ### Sorting 13 notes into 3 categories (11.2s) I asked it to organize notes shuffled in chronological order into categories. The result came out well in table form and the explanation was accurate. However, **there were two problems.** First, the same note ("take the medicine — tomorrow", a typo-laden note) appeared **in two categories at once**, so the total of 13 grew to 14. No deduplication occurred. Second, it **preserved as-is** the typos mixed into the source text. These are the kind of text a person would fix immediately if typing it themselves. ### Sorting a schedule (2.2s) I gave it the condition "only pick the ones with a deadline," but it **included rather than skipped** three items with unclear deadlines. It did mark those items as "needs confirmation," though. The wording was on the generous side, but **failing to recognize the "only" condition** is a clear deviation from the instruction. ### Summarizing long meeting minutes (3.0s) It **accurately extracted four pieces of information**: the release schedule, three QA blockers, the decision to hand off the Windows item, and the marketing draft deadline. This it definitely does well. ### Shell commands — all three actually worked (11.8s) This is the most important verification. I **ran the commands the model proposed on my actual computer** to check. Because command syntax differs by operating system, I'll explain here what tasks I gave it. | Task given | What was verified | |---|---| | Find the top 10 files over 100MB in the home folder, sorted by size | It accurately found a 6.5GB DB file | | Find only files with the `.txt` extension in the Downloads folder | Worked as intended | | Count how many files in the log folder were modified within the last 3 days | Returned "32" | When proposing the third, the model inserted a nonexistent option ahead of the "modified within the last 3 days" condition, so the actual execution failed. **The condition itself was accurate**, and it was the kind of mistake a person could fix immediately with a search. This one small typo was a good example that became the basis for judging that "this model's command knowledge is trustworthy." --- ## Area 2. The ability to choose tools on its own — 60% accuracy I defined four tools for the model (web search / shell execution / file reading / note saving) and instructed it to "answer using these tools." The tool-call arguments were **all accurate**. The problem was the **step of choosing** the tool. | Situation | Expected | Actual | Result | |---|---|---|---| | "What's the weather in Seoul today" | Web search | Even put in the exact search query | Correct | | "Check whether the resume file exists" | Shell execution | Accurately produced a command to check file existence | Correct | | "Save the note 'dentist tomorrow at 3'" | Note saving | Called normally | Correct | | "How full is the disk" | Shell execution | Asked back for the OS, answered directly | Failed | | "Read the resume and suggest improvements" | File reading | "Please upload it," answered directly | Failed | **3 out of 5, 60% accuracy.** The two failures are different in nature. Asking back for the OS on "how full is the disk" is **the most common trap.** The model did not understand that the user is speaking from in front of their own computer. This will always happen unless the agent prompt explicitly states the context "currently running on the local machine." "Read the resume and suggest improvements" means it has no confidence in the file system. It **did not recognize up front** that it could read files itself — even though a file tool had been given to it. --- ## Area 3. Multi-step tasks — this is where it collapses The previous two areas were "one request → one answer." For a personal assistant, the real value comes from **tasks that run three or more steps.** And here the results changed completely. ### Scenario 1: 5 receipts → totals per merchant → save a note I tested with five real receipt files I prepared myself. **Step 1 — success.** The model built a command to read and print every file in the receipts folder one by one, actually ran it, and the contents of all five files were returned as-is. **Step 2 — failure.** The model's reply after receiving the five results was this. > "If you tell me what you'd like done with this information, I'll organize it accordingly. (e.g., calculate the total, sort by date, categorize, etc.)" **It completely forgot the original request.** It neither calculated the total nor saved a note. When tool results come back into the context, at that point it loses track of what it was doing. ### Scenario 2: just calculate the total I made it simpler and just said, "Tell me the total of the receipt amounts." What the model passed to `run_command` was **not a runnable command but annotated pseudocode.** ``` # Need a script that extracts amounts from every file in the receipts folder and calculates the total. # Need to know the actual file structure and how amounts are extracted (e.g. text file, PDF, image). # ... ``` Throwing that into the shell naturally fails. And next it asked back, "Which original request are you referring to?" ### Scenario 3: when the command was specified directly Here performance jumps sharply. | Request | Result | |---|---| | Run "add up all the receipt amounts" | **Total of 111,900 KRW, accurate** (0.4s) | | Count the number in the file-listing output | "File count: 8" — actually 5 (it counted the total line at the top of the output itself and two directory entries as files) | | Sort in reverse and save to a new file | The command succeeded, but it **skipped the verification step of reopening the saved result**, and the final answer was the single line "I've finished your request." | What the second and third share is **the absence of self-verification.** It did not notice the wrong number (8) and reported it as-is, and it quietly skipped the second step of a two-step task. ### The key point — there are only two kinds of failure Categorize this model's failures and there are exactly two. **First, absence of self-verification.** There is no doubt at all about wrong output. If the execution result says "file count: 8," it believes that is the correct answer. **Second, context collapse.** When tool results return, it forgets the original goal. This is not a parameter problem but a structural one. In the end, the usable pattern becomes clear. > **If you specify down to the command, it executes accurately. If you state only the goal, it wanders midway.** --- ## Area 4. Summarizing search results — it does well A test of passing search-engine results to the model to process. **Weather summary + judgment** (0.9s): It accurately summarized the information — a 70% chance of precipitation and rain starting at 3 PM — in one sentence, and even completed the judgment, "You might want to bring an umbrella." **Organizing interest-rate figures** (2.1s): It accurately organized the policy rate of 2.25%, lending rates of 3.4–4.6%, and leasehold deposit loans of 3.9–4.3% into a table. It did not go outside the given range. It does well at **summarizing the search results.** The problem, as seen in Area 2, is **deciding on its own whether to run a search.** --- ## How Should It Be Configured ### 1. Increase context length gradually Measured VRAM values. | num_ctx | VRAM | Headroom left on 8GB | |---|---|---| | 8,192 | 4,039 MiB | 4,153 MiB | | 32,768 | 4,372 MiB | 3,820 MiB | | 65,500 | 4,801 MiB | 3,391 MiB | Even increasing the context 8× from 8k to 64k grows VRAM by only **about 760 MiB**. This is because the gemma family uses sliding-window attention, so the cache of the later layers barely grows. **So there is no problem raising it to 64k from the start.** If anything, insufficient context truncates search results and long documents, making performance drop sharply. However, **when about 1GB of headroom remains, that is the right level.** If you want to load two models at once, it is safer to start at 32k. ### 2. To increase multi-turn stability The most effective setting is **enabling thinking mode.** This test had it disabled, but the failure types in Area 3 (losing the goal, returning meaningless questions) are largely alleviated when thinking is on. On the other hand, **speed drops significantly** and thinking proceeds in English, so it is better to turn it on in long agent loops and off for simple queries. Recommended parameters: ``` # for simple queries and summaries (fast, thinking off) PARAMETER num_ctx 65500 PARAMETER temperature 0.4 PARAMETER top_p 0.9 PARAMETER top_k 20 PARAMETER num_predict 3072 ``` ### 3. Sentences that must go into the system prompt The two failures in Area 2 can be fixed with the prompt. ``` - You are running locally on the user's computer. - Do not ask the user which OS. You already know which OS you are running on. - You can read files directly. If a file tool is provided, do not ask for an upload. - When you receive a tool result, always recall the original request first. - If the result differs from expectations, do not just report it — check the cause. ``` The fourth sentence is especially important, because the failure that appeared repeatedly in this test was exactly "receiving results and forgetting the original goal." ### 4. Memory optimization settings In an 8GB environment, you only need to add the two lines below. ```bash OLLAMA_KV_CACHE_TYPE=q8_0 # KV cache in 8bit -> about 50% memory savings OLLAMA_FLASH_ATTENTION=1 ``` Previously there was a problem where the model would not load at 64k. **After applying `OLLAMA_KV_CACHE_TYPE=q8_0`, it loads normally at 4,815 MiB.** If you are trying local AI in an 8GB environment, this variable is essential. ### 5. Model-switching cost | Situation | Time taken | |---|---| | First request after unloading (cold load) | 3.10s | | Consecutive requests while resident | 0.13s | | Switching from another model | about 6–7s | With `OLLAMA_KEEP_ALIVE=5m`, a model unused for 5 minutes is automatically unloaded. If you plan to alternate between two models, you have to accept **"waiting a few seconds on every switch."** --- ## Measured speed comparison of the two models Compared with the same prompt (polishing "today's meeting ran 3 hours with no conclusion"). | Model | Speed | Output | |---|---|---| | supergemma-e4b-q4km-63k | 75.5 tok/s | Today's meeting ran three whole hours and still reached no conclusion. | | qwen3.8-4b-q6k | 85.5 tok/s | Even after a meeting that ran a full three hours today, no conclusion was reached. | Neither model could stay resident **simultaneously** at 64k on 8GB. (4,815 + 5,000 = 9,815 MiB > 8,192 MiB) A strategy of using both models with manual switching is realistic. --- ## So Is This Model Usable **Breaking it down by area.** **Areas you can use right away** - Long-document summarization, meeting-minute organization - Sentence polishing, email draft writing - Shell-command knowledge queries - Search-result summarization **Areas requiring human oversight** - Judgment in choosing tools on its own (60%) - Tasks that run three or more steps - Calculations that need number verification **Areas you should not use it for** - Automated execution involving important data - Judgment requiring self-verification ## A Question That Remains at the End In this test, the total of the five receipts I made is **111,900 KRW.** The model produced the exact number when given a single command. But when the same task was stated only as a "goal," it got lost midway. I felt that **we have reached a point where the power to hold onto a goal to the end matters more than the model's intelligence.** And that this power must be reinforced through prompts and design, not parameters. So in the next record I plan to compare before and after prompt reinforcement. Same scenario, different instruction style. I'll check how much it improves and whether that is worth recording.