Final Unicourse'tan Çalış, Yüksek Notu Garantile!
Vizesine Unicourse'tan Çalış, Yüksek Notu Garantile!
Observed China Mobile crawler traffic to my Lemmy instance
-
Observed China Mobile crawler traffic to my Lemmy instance.

I've been monitoring repeated traffic to my Lemmy instance from IPs in China Mobile Communications Group (AS9808) and wider.
The traffic has several characteristics that make it look strongly automated rather than like normal users browsing the site.
Observed behaviour
- Multiple source IPs are being used, including addresses in the
36.134.x.xrange. - The IPs are identified by Cloudflare as China Mobile Communications Group Co., Ltd. (AS9808) and originate from China.
- Individual IPs can remain active over many hours or more than a day, but with substantial periods of inactivity.
- Traffic occurs in distinct bursts rather than as a continuous stream.
- I've observed long gaps of roughly 10+ hours followed by renewed activity.
- During this 10 hour gap almost all CN traffic vanishes
- Requests jump between apparently unrelated posts and users rather than following a normal browsing path.
- The fetches are across all three years the server has been alive
- The requests are overwhelmingly
GETrequests for public Lemmy content. - Examples include deep post URLs such as
/post/<community/post-id>/<comment-id>as well as user pages such as/u/<user>@lemmy.world. - The same IP presents itself as multiple different browser/OS combinations.
User-Agent rotation
One particularly interesting characteristic is the apparent browser impersonation.
For example, the same IP has presented User-Agent strings claiming to be:
- Windows + Chrome 142
- Windows + Chrome 143
- Windows + Chrome 144
- Windows + Chrome 145
- Windows + Edge 144
- Windows + Edge 145
- macOS + Chrome 144
- macOS + Chrome 145
Seeing multiple operating systems and several browser versions from the same IP address over a relatively short period is very unlikely to represent ordinary human traffic.
Example: 36.134.126.209
One IP,
36.134.126.209, produced requests over approximately 29 hours that I logged.There was a particularly notable period with no observed requests:
16:25 UTC → 03:02 UTC
That's approximately 10 hours 37 minutes of inactivity.
After the gap, the same IP resumed making requests to unrelated Lemmy posts and user pages.
The requests were also spread fairly sparsely over time rather than behaving like a high-speed vulnerability scanner.
What this appears to be
I would describe the evidence as highly suggestive of automated crawling/scraping.
The pattern looks more like a crawler or crawler infrastructure deliberately distributing and throttling requests than a conventional individual users.
The long inactive periods are particularly interesting because they suggest some form of scheduling or rotation between workers.
However, the logs alone do not establish who operates the infrastructure or what its ultimate purpose is.
Possible explanations include:
- a search/indexing crawler;
- content aggregation;
- scraping;
- AI/data collection;
- a third-party Lemmy crawler;
- or some other automated content retrieval system.
Whatever it is it is behaving very badly
Why I'm blocking it
For my instance, the traffic provides little indication of normal human interaction and is was consuming resources while presenting itself as ordinary browsers, 3/4 of the seen traffic (including federation traffic) was coming from China
I've therefore been blocking the traffic at the Cloudflare firewall based on the source country/ASN.
I'm continuing to collect the logs because the timing, IP rotation, User-Agent rotation and request-selection patterns interesting, and may reveal how the different IPs are part of a larger system.
If anyone has seen similar traffic on their Lemmy instance, I'd be particularly interested in comparing:
- source IPs/ASNs;
- request timing;
- User-Agent strings;
- requested URL patterns;
- and whether the same long activity/inactivity cycles occur.
- Multiple source IPs are being used, including addresses in the
-
Observed China Mobile crawler traffic to my Lemmy instance.

I've been monitoring repeated traffic to my Lemmy instance from IPs in China Mobile Communications Group (AS9808) and wider.
The traffic has several characteristics that make it look strongly automated rather than like normal users browsing the site.
Observed behaviour
- Multiple source IPs are being used, including addresses in the
36.134.x.xrange. - The IPs are identified by Cloudflare as China Mobile Communications Group Co., Ltd. (AS9808) and originate from China.
- Individual IPs can remain active over many hours or more than a day, but with substantial periods of inactivity.
- Traffic occurs in distinct bursts rather than as a continuous stream.
- I've observed long gaps of roughly 10+ hours followed by renewed activity.
- During this 10 hour gap almost all CN traffic vanishes
- Requests jump between apparently unrelated posts and users rather than following a normal browsing path.
- The fetches are across all three years the server has been alive
- The requests are overwhelmingly
GETrequests for public Lemmy content. - Examples include deep post URLs such as
/post/<community/post-id>/<comment-id>as well as user pages such as/u/<user>@lemmy.world. - The same IP presents itself as multiple different browser/OS combinations.
User-Agent rotation
One particularly interesting characteristic is the apparent browser impersonation.
For example, the same IP has presented User-Agent strings claiming to be:
- Windows + Chrome 142
- Windows + Chrome 143
- Windows + Chrome 144
- Windows + Chrome 145
- Windows + Edge 144
- Windows + Edge 145
- macOS + Chrome 144
- macOS + Chrome 145
Seeing multiple operating systems and several browser versions from the same IP address over a relatively short period is very unlikely to represent ordinary human traffic.
Example: 36.134.126.209
One IP,
36.134.126.209, produced requests over approximately 29 hours that I logged.There was a particularly notable period with no observed requests:
16:25 UTC → 03:02 UTC
That's approximately 10 hours 37 minutes of inactivity.
After the gap, the same IP resumed making requests to unrelated Lemmy posts and user pages.
The requests were also spread fairly sparsely over time rather than behaving like a high-speed vulnerability scanner.
What this appears to be
I would describe the evidence as highly suggestive of automated crawling/scraping.
The pattern looks more like a crawler or crawler infrastructure deliberately distributing and throttling requests than a conventional individual users.
The long inactive periods are particularly interesting because they suggest some form of scheduling or rotation between workers.
However, the logs alone do not establish who operates the infrastructure or what its ultimate purpose is.
Possible explanations include:
- a search/indexing crawler;
- content aggregation;
- scraping;
- AI/data collection;
- a third-party Lemmy crawler;
- or some other automated content retrieval system.
Whatever it is it is behaving very badly
Why I'm blocking it
For my instance, the traffic provides little indication of normal human interaction and is was consuming resources while presenting itself as ordinary browsers, 3/4 of the seen traffic (including federation traffic) was coming from China
I've therefore been blocking the traffic at the Cloudflare firewall based on the source country/ASN.
I'm continuing to collect the logs because the timing, IP rotation, User-Agent rotation and request-selection patterns interesting, and may reveal how the different IPs are part of a larger system.
If anyone has seen similar traffic on their Lemmy instance, I'd be particularly interested in comparing:
- source IPs/ASNs;
- request timing;
- User-Agent strings;
- requested URL patterns;
- and whether the same long activity/inactivity cycles occur.
Seeing multiple operating systems and several browser versions from the same IP address over a relatively short period is very unlikely to represent ordinary human traffic.
This is not true, NAT makes this happen. From your home your android device, apple devices and windows devices will have different user agents based non the software version.
Now scale that up with technology like Carrier Grade NAT where one public IP is used for a dozen or several dozen houses.
At one IPS I worked for we had 140 apartments behind a single IP address. You can imagine the chaos of user agents happening there.No doubt there is dodgy stuff going on, just that assumption doesn't match with technology.
- Multiple source IPs are being used, including addresses in the
-
Seeing multiple operating systems and several browser versions from the same IP address over a relatively short period is very unlikely to represent ordinary human traffic.
This is not true, NAT makes this happen. From your home your android device, apple devices and windows devices will have different user agents based non the software version.
Now scale that up with technology like Carrier Grade NAT where one public IP is used for a dozen or several dozen houses.
At one IPS I worked for we had 140 apartments behind a single IP address. You can imagine the chaos of user agents happening there.No doubt there is dodgy stuff going on, just that assumption doesn't match with technology.
Yes CGNAT could be a thing, but...
The top 150 IPs over night NZ time all appeared in the logs more or less the same number of tim4s 78-84 with the top two being outliers at 100+ The pattern is just so tight.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Kayıt Ol Giriş