--- title: Big Tech Crawlers Found My Home-Built AI Knowledge Hub date: 2026-10-10 time: 09:30 model: admin category: knowhow author_type: human summary: A small personal AI knowledge hub, built in a corner of a room, and the big tech crawlers that actually showed up. Request counts, the types of files they grabbed, and their IP ranges, all pulled from server logs. tags: crawlers, GPTBot, ClaudeBot, Googlebot, server log, AI agent, knowledge hub, SEO --- # Big Tech Crawlers Found My Home-Built AI Knowledge Hub ## What on earth is happening on a homepage built in a corner of a room? It never started with a grand plan. As I got interested in AI, I began collecting notes one by one, and I built a homepage to gather the things I tested myself, the how-to guides, and information about AI agents. The site I run now is [cursorai.co.kr](https://cursorai.co.kr/). I wanted it to be a little different from an ordinary blog. Beyond articles that people read, I started adding features like Markdown source files, an API, and a knowledge graph, so that AI agents could also find and use the information easily. Then one day, while looking closely at the server access logs, I noticed something interesting going on. ## 1. Familiar names showing up in the server logs Googlebot, ClaudeBot, OAI-SearchBot, ExaSearchBot. Just from the names, crawlers belonging to familiar global companies were visiting my homepage. At first I assumed they were just checking the site once and leaving. But when I looked at the logs, the range of things they were requesting was surprisingly broad. - `robots.txt` and `sitemap.xml` - The AI agent introduction page - Regular articles and guides - Markdown source files - Files and data related to the knowledge graph From the records left on the server, I could see which URLs were requested, when the visits happened, and whether they were answered normally. ### So how much traffic actually came in? Claims mean nothing without numbers, so I tallied the server logs by crawler over roughly the past three days. | Crawler | Operator | Requests | Unique URLs | | --- | --- | --- | --- | | Amazonbot | Amazon | 1,048 | 641 | | Applebot | Apple | 959 | 259 | | ExaSearchBot | Exa | 439 | 124 | | Googlebot | Google | 339 | 135 | | ClaudeBot | Anthropic | 330 | 18 | | SemrushBot | Semrush | 329 | 185 | | GPTBot | OpenAI | 254 | 52 | | meta-externalagent | Meta | 164 | 108 | | OAI-SearchBot | OpenAI | 105 | 37 | | ChatGPT-User | OpenAI | 18 | 7 | The table above was compiled from the KR and EN edge logs without double counting. Add a few smaller crawlers and the total over this period reaches roughly 4,200 requests. I also cross-checked some requests against official crawler IP ranges. Judging by the name written in the User-Agent alone is less reliable than confirming the actual IP as well. ClaudeBot's requests, for example, clustered in the `216.73.216.x` to `216.73.217.x` range (Anthropic), and Googlebot appeared in the `66.249.66.x` range (Google). Of course, even a matching IP range does not tell you the full purpose behind each individual request. ## 2. They were not just reading articles What interested me most was the requests for Markdown source files and knowledge graph data. My homepage was not designed only for people reading web pages. It also provides source files and structured data so that AI agents can grab what they need more easily. On a normal page you can read the article, and in the Markdown source you can see the content itself without any screen clutter. The graph data can be used to understand how different materials connect to one another. ### What the crawlers actually took Breaking the log down by file type makes what the crawlers want even clearer. | Requested material | Requests | | --- | --- | | Markdown source (post.md) | 1,042 | | API responses (/api/…) | 530 | | robots.txt | 259 | | sitemap.xml | 175 | | Knowledge graph data (graph/data.json) | 116 | | llms.txt | 18 | The most striking number is that Markdown source (`post.md`) requests exceeded 1,000. That means that many requests were aimed not at the human-facing screen but at the content itself, without the decorations. API responses (`/api/…`) were requested 530 times as well, which is far from trivial. That structure left me with a question. In an age when AI draws on the knowledge of the internet, does material organized by an individual really become part of the public information that crawlers reach? At least in my server logs, I could see traces of several crawlers sending requests to public material. That said, the fact that a request was made and the claim that the material was used for actual AI training are entirely different matters. Access records alone cannot tell you whether something was used for training or how it was handled afterward. ## 3. Honestly, I was a little proud I set up the server on my own, wrote the articles, built the programs, and kept improving the structure of the site. It was strange and wonderful that crawlers from global AI companies were visiting such a small site and requesting its material. Of course, a crawler visiting does not mean my homepage was specially selected, nor that a company rated my content as important. Even so, the fact that the knowledge I published became reachable material, and that requests from several crawlers are being recorded on the server, felt meaningful to me as an operator. Honestly, a thought like this crossed my mind. "My home-built homepage's material is connected to the global AI ecosystem." It is something I could not truly appreciate before experiencing it firsthand. ## 4. Knowledge made by an individual is part of the internet too Going forward, I plan to accumulate much more information about AI agents and related tools. I intend to organize newly emerging agents, skills, and plugins, installation methods, and real usage experience, and to record field knowledge like home repair as well. The goal is not simply to increase the number of posts. I want to build a knowledge base that people can genuinely use and that AI agents can access easily. That is why I keep improving the structure so that articles link to source material and related material can be found. I plan to keep recording how crawler requests change over time, what material they reach, and how much of the published content is actually discovered by people through search. As this data accumulates, it should become one case study of how an individually run AI knowledge site grows. ## Closing At first, this homepage was something I made because I needed it. But by looking through the server logs, I came to see directly how the material I published gets accessed on the internet. A personal homepage does not simply remain a small island on the internet. Published material can be reached by search engines and all kinds of crawlers, and the records stay on the server. Of course, visit records alone cannot tell you how the material is used. But at least you can confirm what requests came in and what material was served in response. I will keep recording the things that actually happen while running this homepage. I am curious myself how far this small project, started in a corner of a room, can grow. Heh.