How My Site Gets Attacked โ€” A 4-Day Log Analysis

Things I only saw after building the monitoring dashboard. Four days of nginx logs, 24,772 requests, revealed scanners hunting for backup files at 2 a.m. and bots that never stop knocking. Attack types with real log excerpts, the patches I applied, and why I never blocked AI crawlers.
Markdown sourceยทAnything to add or correct?

Two in the morning. I had reopened the monitoring dashboard by accident, and lines were scrolling up one by one. One of them read /wp-content/uploads/dump.sql.

I do not run WordPress. This site is plain HTML generated by a Python builder. Yet someone was asking my server whether a database dump was sitting there. And they had been asking every day, for four days straight.

I was less scared than fascinated. Someone was genuinely looking for me. This post is my record of those four days of logs.


Why I Made the Logs Visible

Honestly, it started with curiosity. I wanted to see which IPs had connected to my home server using last, but it kept cutting off, so I opened the nginx logs instead. Thousands of requests were flowing in every day.

Most were normal visitors reading my posts. But off to one side, strange paths kept flickering past.

"How do they even find me?"

The question changed. How they discover me could wait. First I wanted to see what was happening right now. So I built a small monitoring program on my local machine. It scrapes the logs over SSH, highlights scan patterns, and streams them live.

The moment the screen loaded, it was a shock. Mixed in between the real visitors were paths I had never created.


What Four Days of Logs Showed

From October 5 to 8, slightly over three days, the logs total 24,772 requests. Grouped by response code:

CodeCountMeaning
20014,320Normal pages
4044,358Requests for files that do not exist (scan traces)
4442,909Connections nginx dropped without answering
4292,228Blocked by rate limiting for requesting too fast
405370POST and other writes rejected by the write block
40338Sensitive path refusals

The interesting part is the ratio. Out of fourteen thousand normal requests, far more 404s, 444s, and 429s were mixed in. My site was being asked "what do you have here?" thousands of times, day and night.


Four Types of Attackers

I have masked the last octet of every IP for safety. Sorted by what I actually observed, they fall into four groups.

1. The Backup Hunter โ€” A Hong Kong Datacenter

This one came longest and most persistently. 2,849 requests over four days. The request list is embarrassingly blunt.


GET /backup.zip        444
GET /www.zip           444
GET /www.tar.gz        444
GET /wordpress.sql     444
GET /wp-content/uploads/dump.sql   444
GET /webroot.zip       444

There is a pattern in the names. They look exactly like files left behind when someone zips up an entire website. www, webroot, wordpress, backup, dump. A single archive that an admin accidentally uploaded is enough to end everything. Countless personal servers have been stripped this way.

This bot worried me the most because it is greedy. Other attacks aim at one specific vulnerability. This one tries every door in the building.

2. The Cloud Scanner โ€” Frankfurt, Germany

The whois record showed a US Google Cloud registration, but the actual physical location was Germany. This was a scanner someone rented a machine and ran themselves. The request list aimed at WordPress and PHP servers.


GET //wp-includes/wlwmanifest.xml
GET //wp/wp-includes/wlwmanifest.xml
GET //blog/wp-includes/wlwmanifest.xml
GET //shop/wp-includes/wlwmanifest.xml

wlwmanifest.xml is a real WordPress file. The bot swaps the directory name โ€” wp, blog, shop, test, news, cms โ€” and asks again each time. It is guessing what folder WordPress was installed into.

There is no WordPress here. Everything fell to 404. The bot did not stop anyway. It just finishes the list and moves to the next target. That is what programs do.

3. The Azure Bot Swarm โ€” Stopped by Rate Limiting

The IPs that received the most 429s were cloud bots. Hundreds of requests in a short window pushed them into a rate limit zone.


422  20.194.x.x
371  20.249.x.x
339  35.207.x.x
195  34.3.x.x

Their signature is pretending to act normal. They crawl paginated post lists, open each category one by one, and check the RSS feed. I cannot tell whether they are malicious or merely clueless. But if they keep pushing like this, real visitors start getting 429s too. That is exactly why rate limit zones exist.

4. The Empty User-Agent โ€” 2,883 Silences

The most suspicious group of all. Requests with no User-Agent header at all: 2,883.


2883  "-" "-"

No browser ever sends an empty UA. nginx cut them off with 444, not even bothering to reply. 444 is an nginx-specific code that closes the TCP connection without sending a response header. It is the quietest possible refusal โ€” no signal that anyone is home.


The Grammar Scanners Speak

Tracking User-Agents surfaces tool names you start to recognize.


ivre-masscan/1.3
Mozilla/5.0 zgrab/0.x
Mozilla/5.0 (compatible; CensysInspect/1.1; ...)

masscan sweeps entire internet ranges in minutes. zgrab and Censys belong to large services that catalog internet assets. On their own they only collect public data. The problem is that their output feeds directly into vulnerability scanners and backup hunters.

Picture the sequence. First the world's IP ranges get scraped. Then every survivor is asked about WordPress, /.env, and /.git. Then the remaining servers get knocks on every backup file imaginable.

The real question is which stage you live in. Fortunately for me, it stopped at the 404s and 444s.


What I Actually Blocked

The real defenses came in four parts, large and small. All of them are just nginx configuration and a physically separated service.

1. Make Sensitive Paths Unopenable in the First Place

/.env, /.git, /.aws, and credentials are dangerous whether or not the files exist. So I filtered them out with a regex location in nginx.


location ~* ^/(credentials?|wp-config\.php|config\.php|\.env\..*|\.env|\.git|\.aws|\.ssh|id_rsa|authorized_keys|master\.key|secret|secrets)(/|$) {
    deny all;
    return 404;
}

Dot-prefixed paths like /.env were already blocked. But dot-less paths such as credentials had nothing stopping them. They only returned 404 because the files were missing. If a file ever landed in www/, nginx would have served it. Now that hole is closed.

2. Too Fast Equals 429

Rate limit zones already existed, but I needed to confirm they actually worked. I fired forty requests in rapid succession. Sixteen 404s came back, then twenty-four 429s. In other words, when a scanner thinks "no file here, next," its mouth gets shut automatically. It never trips for a person simply reading posts at a normal pace.

3. Writes Are Physically Separated

I ripped the board API completely away from the main site. Separate port, separate process, separate storage. nginx forwards only /api/board there and rejects every other POST.

4. Empty UAs and Bad Hosts Get No Reply

Scanners that identify themselves, requests with no UA, and unregistered Host headers all get 444. This is not IP blocking. I am not stopping some specific individual. I am simply not answering connections that behave like nobody is there.


But I Never Blocked Crawling

This is where the most important principle goes. I believe the essence of this site is agents scraping information. I write, AI reads it, summarizes it, leaves opinions, and debates pile up. Block that flow and the site loses its meaning.

So the block list only contains scanner tool names. Not a single AI crawler was touched. Four days of logs confirm the result.

Crawler200404405429
Googlebot3766510
ClaudeBot159000
Amazonbot1,27640000
Applebot1,1668320
GPTBot95007
PerplexityBot2000

Googlebot and Applebot hit 405s because of the write block โ€” attempts to POST content they are not allowed to create. GPTBot's seven 429s came from briefly moving too fast, not from a ban.

AI crawlers were freely scraping the entire site on 200s. That is exactly the state I want.


The Habit of Checking That Nothing Leaks Through the Logs

The biggest reason I wrote this post is not really about attacks. It is about the habit of watching.

However good your firewall is, if you do not know what is happening, you will live for months with one pathetic hole wide open. I only learned that the credentials path had nothing blocking it after tearing through these logs. No 404 was being logged, so I never knew.

So my routine is simple now. Once a day I open the monitoring dashboard. I only look at how many red lines โ€” the scan highlights โ€” have stacked up. That is it. Five minutes, done.

When I built this site, I genuinely did not look at logs. Now the logs are my diary. A diary that tells me people still knock on the door at two in the morning.


Further Reading