<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Observed China Mobile crawler traffic to my Lemmy instance]]></title><description><![CDATA[<h1>Observed China Mobile crawler traffic to my Lemmy instance.</h1>
<p dir="auto"><img src="https://no.lastname.nz/pictrs/image/4b9bcd7e-f742-4357-b8e9-7c08d276babd.png" alt="" class=" img-fluid img-markdown" /></p>
<p dir="auto">I've been monitoring repeated traffic to my Lemmy instance from IPs in <strong>China Mobile Communications Group (AS9808)</strong> and wider.</p>
<p dir="auto">The traffic has several characteristics that make it look strongly automated rather than like normal users browsing the site.</p>
<h2>Observed behaviour</h2>
<ul>
<li>Multiple source IPs are being used, including addresses in the <code>36.134.x.x</code> range.</li>
<li>The IPs are identified by Cloudflare as <strong>China Mobile Communications Group Co., Ltd. (AS9808)</strong> and originate from China.</li>
<li>Individual IPs can remain active over many hours or more than a day, but with substantial periods of inactivity.</li>
<li>Traffic occurs in distinct bursts rather than as a continuous stream.</li>
<li>I've observed <strong>long gaps of roughly 10+ hours</strong> followed by renewed activity.</li>
<li>During this 10 hour gap <strong>almost all CN traffic vanishes</strong></li>
<li>Requests jump between apparently unrelated posts and users rather than following a normal browsing path.</li>
<li>The fetches are across all <strong>three years</strong> the server has been alive</li>
<li>The requests are overwhelmingly <code>GET</code> requests for public Lemmy content.</li>
<li>Examples include deep post URLs such as <code>/post/&lt;community/post-id&gt;/&lt;comment-id&gt;</code> as well as user pages such as <code>/u/&lt;user&gt;@lemmy.world</code>.</li>
<li>The same IP presents itself as multiple different browser/OS combinations.</li>
</ul>
<h2>User-Agent rotation</h2>
<p dir="auto">One particularly interesting characteristic is the apparent browser impersonation.</p>
<p dir="auto">For example, the same IP has presented User-Agent strings claiming to be:</p>
<ul>
<li>Windows + Chrome 142</li>
<li>Windows + Chrome 143</li>
<li>Windows + Chrome 144</li>
<li>Windows + Chrome 145</li>
<li>Windows + Edge 144</li>
<li>Windows + Edge 145</li>
<li>macOS + Chrome 144</li>
<li>macOS + Chrome 145</li>
</ul>
<p dir="auto">Seeing multiple operating systems and several browser versions from the <strong>same IP address</strong> over a relatively short period is very unlikely to represent ordinary human traffic.</p>
<h2>Example: 36.134.126.209</h2>
<p dir="auto">One IP, <code>36.134.126.209</code>, produced requests over approximately <strong>29 hours</strong> that I logged.</p>
<p dir="auto">There was a particularly notable period with no observed requests:</p>
<p dir="auto"><strong>16:25 UTC → 03:02 UTC</strong></p>
<p dir="auto">That's approximately <strong>10 hours 37 minutes</strong> of inactivity.</p>
<p dir="auto">After the gap, the same IP resumed making requests to unrelated Lemmy posts and user pages.</p>
<p dir="auto">The requests were also spread fairly sparsely over time rather than behaving like a high-speed vulnerability scanner.</p>
<h2>What this appears to be</h2>
<p dir="auto">I would describe the evidence as <strong>highly suggestive of automated crawling/scraping</strong>.</p>
<p dir="auto">The pattern looks more like a crawler or crawler infrastructure deliberately distributing and throttling requests than a conventional individual users.</p>
<p dir="auto">The long inactive periods are particularly interesting because they suggest some form of scheduling or rotation between workers.</p>
<p dir="auto">However, the logs alone <strong>do not establish who operates the infrastructure or what its ultimate purpose is</strong>.</p>
<p dir="auto">Possible explanations include:</p>
<ul>
<li>a search/indexing crawler;</li>
<li>content aggregation;</li>
<li>scraping;</li>
<li>AI/data collection;</li>
<li>a third-party Lemmy crawler;</li>
<li>or some other automated content retrieval system.</li>
</ul>
<p dir="auto">Whatever it is it is behaving very badly</p>
<h2>Why I'm blocking it</h2>
<p dir="auto">For my instance, the traffic provides little indication of normal human interaction and is was consuming resources while presenting itself as ordinary browsers, <strong>3/4 of the seen traffic</strong> (including federation traffic) was coming from China</p>
<p dir="auto">I've therefore been blocking the traffic at the Cloudflare firewall based on the source country/ASN.</p>
<p dir="auto">I'm continuing to collect the logs because the <strong>timing, IP rotation, User-Agent rotation and request-selection patterns</strong> interesting, and may reveal how the different IPs are part of a larger system.</p>
<p dir="auto">If anyone has seen similar traffic on their Lemmy instance, I'd be particularly interested in comparing:</p>
<ul>
<li>source IPs/ASNs;</li>
<li>request timing;</li>
<li>User-Agent strings;</li>
<li>requested URL patterns;</li>
<li>and whether the same long activity/inactivity cycles occur.</li>
</ul>
]]></description><link>https://forum.ieu.app/topic/a498d5ce-1d13-48d6-962e-e05f2509cbaf/observed-china-mobile-crawler-traffic-to-my-lemmy-instance</link><generator>RSS for Node</generator><lastBuildDate>Sun, 06 Sep 2026 04:28:32 GMT</lastBuildDate><atom:link href="https://forum.ieu.app/topic/a498d5ce-1d13-48d6-962e-e05f2509cbaf.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 17 Aug 2026 08:23:58 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Observed China Mobile crawler traffic to my Lemmy instance on Tue, 18 Aug 2026 01:40:43 GMT]]></title><description><![CDATA[<p dir="auto">Yes CGNAT could be a thing, but...</p>
<p dir="auto">The top 150 IPs over night NZ time all appeared in the logs more or less the same number of tim4s 78-84 with the top two being outliers at 100+ The pattern is just so tight.</p>
]]></description><link>https://forum.ieu.app/post/https://no.lastname.nz/comment/7071370</link><guid isPermaLink="true">https://forum.ieu.app/post/https://no.lastname.nz/comment/7071370</guid><dc:creator><![CDATA[blueether@no.lastname.nz]]></dc:creator><pubDate>Tue, 18 Aug 2026 01:40:43 GMT</pubDate></item><item><title><![CDATA[Reply to Observed China Mobile crawler traffic to my Lemmy instance on Mon, 17 Aug 2026 09:39:13 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto">Seeing multiple operating systems and several browser versions from the same IP address over a relatively short period is very unlikely to represent ordinary human traffic.</p>
</blockquote>
<p dir="auto">This is not true, NAT makes this happen.  From your home your android device, apple devices and windows devices will have different user agents based non the software version.</p>
<p dir="auto">Now scale that up with technology like Carrier Grade NAT where one public IP is used for a dozen or several dozen houses.<br />
At one IPS I worked for we had 140 apartments behind a single IP address.  You can imagine the chaos of user agents happening there.</p>
<p dir="auto">No doubt there is dodgy stuff going on, just that assumption doesn't match with technology.</p>
]]></description><link>https://forum.ieu.app/post/https://lemmy.world/comment/25338382</link><guid isPermaLink="true">https://forum.ieu.app/post/https://lemmy.world/comment/25338382</guid><dc:creator><![CDATA[slazer2au@lemmy.world]]></dc:creator><pubDate>Mon, 17 Aug 2026 09:39:13 GMT</pubDate></item></channel></rss>