Residential Proxies, Explained From the Detection Side

A residential proxy routes traffic through an IP address an ISP assigned to a real household or small-business connection. The destination site sees a Comcast or Vodafone address instead of an AWS one, so every control that asks "is this a datacenter IP" comes back clean. That is the whole product. Rotation, geo-targeting, and sticky sessions are packaging around that one property.

Almost every page on this topic is written by a company that sells proxies, so the detection question gets one line and no mechanism. This one answers it. If you run signup or login on a product with open registration, this traffic is already in your logs, and your IP controls are not missing it because they are badly tuned.

How does a residential proxy work?

You never connect to a specific household. You connect to a back-connect gateway run by the provider, using credentials that also encode your targeting parameters. The gateway picks an exit node matching those parameters (country, region, city, ASN, session length), tunnels your request through that device, and routes the response back. The exit node is software running on a consumer laptop, phone, router, or streaming box, and it exists only while that device is online.

Three delivery models sit on top of that architecture.

Mobile proxies are a fourth category, exiting through cellular-carrier addresses. They are the hardest tier to act on because of carrier-grade NAT: thousands of real subscribers share one public address, so blocking it has a blast radius you cannot see.

Residential vs datacenter vs mobile proxies

DatacenterResidentialMobile
IP registered toCloud or hosting providerConsumer ISPCellular carrier
What an ASN lookup returnsA hosting organizationA consumer ISP, same as your real usersA carrier, same as your real mobile users
Shared with real usersNoYes, the device owner is usually online at the same timeYes, many subscribers share one address via CGNAT
Session stabilityHighVariable, the node disappears when the device doesVariable
Does IP-layer blocking workYes, the ASN gives it awayNo, the address is indistinguishable from a subscriber'sNo, and one block can take out thousands of real subscribers

Where do the IP addresses come from?

Two ways. The FBI's Cyber Division put it plainly in a March 2026 public service announcement: "Residential proxies obtain residential IP addresses from devices in two ways: The owner of the device provides consent, or the owner of the device does not provide consent and is unaware their IP address is being used."

The same PSA names five acquisition vectors:

  1. SDK partnerships. Proxy services pay app developers to embed a bandwidth-sharing SDK. Users install the app, accept the terms, and the SDK routes proxy traffic in the background.
  2. Free VPNs with hidden terms of service. Free VPN apps enroll their own users' devices, with the disclosure buried in terms nobody reads or written in language nobody parses.
  3. Compromised IoT devices. Home networks reached through devices configured with malicious software before purchase, or backdoored during a later app download.
  4. Malware. Free game content, pirated software, and torrented media carrying a payload that enlists the device as an exit node.
  5. Passive income schemes. Apps that pay users for bandwidth without disclosing what the bandwidth does.

The difference between a consenting node and a compromised one is invisible from the receiving end. The packet is identical.

The documented cases

In January 2026 Google Threat Intelligence Group published an investigation into a network it tracks as IPIDEA. Its finding on vendor marketing: "While many residential proxy providers state that they source their IP addresses ethically, our analysis shows these claims are often incorrect or overstated. Many of the malicious applications we analyzed in our investigation did not disclose that they enrolled devices into the IPIDEA proxy network." Google identified over 600 applications, largely benign in function, carrying monetization SDKs that enabled proxy behavior, and reported that thirteen separately branded proxy and VPN services were all controlled by the actors behind IPIDEA. In a single seven-day period that month it observed over 550 threat groups it tracks using IPIDEA exit nodes, for access to victim SaaS environments, on-premises infrastructure, and password spray attacks.

The starkest case is older. In May 2024 the Department of Justice announced the takedown of the 911 S5 botnet and the arrest of its administrator. The unsealed indictment alleges he spread malware through free VPN programs he operated (MaskVPN and DewVPN) and through pay-per-install bundles attached to pirated software, amassing residential Windows machines associated with more than 19 million unique IP addresses, 613,841 of them in the United States, between 2014 and 2022, and took in roughly $99 million selling access. An indictment is an allegation, as the DOJ said in its own announcement.

BADBOX 2.0 shows a third path: preinstallation. The FBI and HUMAN Security described a botnet of more than one million consumer devices, largely low-cost off-brand tablets, TV boxes, and projectors, many carrying the malware before the customer opened the box. They served as proxy exit nodes, clicked ads in the background, and ran credential stuffing.

Is any of it consensual?

Some of it. Bright Data's SDK license requires partner apps to disclose peer status and offer an opt-out, and other providers state they source exclusively through opt-in SDK integrations. Those models exist, and none has been independently audited.

The one time researchers looked hard, the picture was mixed. A peer-reviewed measurement study presented at IEEE Security and Privacy in 2019 infiltrated the five leading commercial residential proxy services of the time and reported that none of the five ran a fully consent-based network. It identified 237,029 IoT devices acting as exit nodes, and of 67 distinct proxy programs observed, 50 were flagged as malicious by antivirus tooling. Google's conclusion sets the bar: "any claims of 'ethical sourcing' must be backed by transparent, auditable proof of user consent."

For a defender this resolves to one fact. Consensual and non-consensual exit nodes produce identical traffic. You cannot sort them at request time, so sourcing ethics is a procurement problem for the proxy buyer and irrelevant to your detection logic.

How large are these pools really?

Named providers advertise pools from roughly 50 million to over 400 million addresses: NetNut more than 50 million, Webshare 80 million or more, Decodo 115 million or more, Oxylabs 175 million or more, Bright Data 400 million or more. Every figure is self-reported, with no published methodology and no independent audit. Two of those five publish different numbers for themselves on different pages of their own sites.

Set that against the only independent measurement ever published. The 2019 IEEE study observed 6.2 million unique IP addresses relaying its own infiltration traffic across the era's five leading providers combined, over roughly four to five months of fieldwork, 95.22% of which it classified as genuinely residential. That is a lower bound on observed usage, not a ceiling, and it is seven years old. It is also one to two orders of magnitude below what a single vendor claims today, with no independent recount since to establish whether the gap is real growth, marketing, or a change in what is being counted. Google, with direct visibility into one network's command-and-control infrastructure, declined to publish a total at all, saying only that its action cut the available device pool "by millions."

Pricing is the one checkable number here. Across five vendors' public pricing pages, advertised rates run from about $1.40 to $8.00 per GB, converging toward $2 to $4 per GB at moderate monthly commitments.

How is this traffic abused?

Credential stuffing and account takeover. OWASP defines credential stuffing as the automated injection of stolen username and password pairs into login forms: testing known breach credentials for reuse, not guessing. Accertify, a fraud-decisioning vendor, published an analysis of its own client traffic across retail, airlines, quick-service restaurants, marketplaces, ticketing, and grocery for 2024 and 2025. Individual attacks reached 12 million attempts at 1 to 2% credential success rates, producing as many as 200,000 compromised accounts from one campaign. In one attack on an online marketplace, login volume peaked at 300 transactions per second and used 14,864 unique IPs inside a ten-minute window, arriving from T-Mobile, Cox, Charter, Verizon, and Comcast, with individual addresses almost never repeating.

Geography is part of the attack. Pools can be selected down to state and postal code, and the FBI notes buyers choose a country down to the city. An attacker holding stolen bank credentials sources an exit node in the victim's own city, which is exactly where a location heuristic returns green.

Multi-accounting and free-trial abuse. The FBI names spam and fake account creation directly. For SaaS with a free tier this is the same mechanic aimed at signup instead of login: one operator, hundreds of accounts, each from a different consumer IP in a plausible region, every per-IP registration limit satisfied. It is the pattern behind most free-tier abuse that survives basic controls.

Ad fraud. BADBOX 2.0 devices loaded and clicked ads in the background while the same devices were resold as exit nodes. Impressions carrying a residential IP in a high-value market pay a premium, which is what makes it profitable.

Scalping. The FBI lists bypassing purchase restrictions as a standard criminal use: buying limited-release tickets, sneakers, and collectible cards en masse for resale. Per-household purchase limits are enforced against an identifier the attacker buys by the gigabyte.

How do you detect residential proxy traffic?

Why IP reputation alone fails

Not because the data is bad. Because the population it scores has moved.

In Accertify's attack dataset, 99.29% of the IP addresses observed in confirmed attacks carried a clean reputation with no prior fraud signal at the time of use. The rest of that dataset shows the shift in progress: high-risk events with low per-IP usage rose 46% year over year and went from 57.5% to 84.1% of all high-risk events, high-risk events tied to heavily reused IPs fell from 32% to 5.6% over thirteen months, and high-risk transactions from hosting environments dropped 61% year over year.

Read together, those numbers do not say blocklists are inaccurate. They say attackers left the infrastructure blocklists cover. A reputation list records addresses that have already misbehaved. A rotating residential pool is engineered so that no address misbehaves enough, or often enough, to earn a listing before it rotates out of your view. The FBI's own business guidance is still "Block IP addresses that are known to be associated with residential proxy networks," which is a fair statement of the default mental model and, alone, a losing position.

What the IP layer still gives you

The cheapest signal you have. It just cannot be the only one.

ASN and organization lookup is the mechanism under "is this a datacenter." An IP maps to an autonomous system number and an owning organization, which tells you whether the address belongs to a hosting company or a consumer ISP. Commercial IP-intelligence datasets go further and treat residential proxy as its own category, separate from hosting ranges and from registered VPN ranges. MaxMind defines it as addresses on a suspected anonymizing network registered under residential ISPs.

Two details matter in implementation. Newer datasets attach a confidence score and a network-last-seen timestamp instead of a binary flag, which is the right shape: an address enrolled last month and released this month is not the same as one relaying traffic yesterday. And coverage lags reality by construction, so treat a hit as strong evidence and a miss as no evidence. Our free VPN and proxy checker shows what this class of data returns for a single address.

TLS fingerprinting with JA4

The proxy rewrites the source address. It does not rewrite the TLS stack behind it.

JA4, the successor to JA3, fingerprints the ClientHello: TLS version, the set and ordering of cipher suites, the extensions, the ALPN. These are implementation choices, and a given TLS library makes them the same way every time. Real Chrome produces a JA4 that a Python requests session or a Go HTTP client cannot produce, whatever User-Agent it sends and whichever residential exit it leaves from. The fingerprint is emitted as a_b_c, so you can match on partial sections and cluster clients sharing a stack even when the full string differs. The wider JA4+ suite extends the idea: JA4S for the server response, JA4H for the HTTP client, JA4T for TCP, JA4L for latency. We break the format down in what JA4 and JA4T fingerprints are.

The property that matters here: JA4 is computed from the handshake, independent of IP and of geography. Rotating through 14,864 residential addresses does not change it once.

HTTP/2 fingerprinting

Once TLS completes, the client opens an HTTP/2 connection, and the opening frames identify it before a single request header is parsed. Akamai's Black Hat EU 2017 research defined a fingerprint from four client-controlled parts of the connection preface: SETTINGS frame parameters, WINDOW_UPDATE behavior, PRIORITY frames, and pseudo-header order.

Each is a fixed decision in the client's HTTP/2 implementation. A scripted client declaring itself Chrome, behind a clean residential IP, still sends a preface Chrome does not send. This layer catches the automation stacks that survive TLS fingerprinting by copying a browser's cipher list without copying its HTTP/2 stack.

Browser and device fingerprinting

The proxy anonymizes the network path. It does nothing to the device.

Browser fingerprinting captures the execution environment: rendering behavior, hardware and API surface, font and codec availability, and whether what the browser claims matches what it demonstrably is. The result is a device-level identifier that survives IP rotation, incognito, and cookie clears.

This is where the attacker's architecture works against them. A rotating pool exists to make one operator look like thousands of users at the IP layer. If the device fingerprint is stable, that operator collapses back into one entity and the pool becomes evidence instead of cover: one device, four hundred addresses, forty ISPs, one hour. Real users do not produce that shape. Trueguard is built on this correlation, terminating sessions against a TLS-terminating fingerprint service so the JA4 and HTTP/2 signals land on the same verification event as the browser fingerprint, the IP intelligence, and the email signal.

Velocity, measured against the right denominator

Per-IP velocity is the control this traffic is built to defeat, and it defeats it by arithmetic. Twelve million attempts spread across a pool large enough that no address crosses your threshold produces zero alerts from a counter that is working perfectly.

Move the denominator off the IP. Count per account, per username across the whole login surface, per device fingerprint, per TLS fingerprint, per ASN, per registration form. In the marketplace attack above, the per-IP view showed nothing while the aggregate showed 300 logins per second. The aggregate was always there. Nobody was counting on an axis the attacker could not rotate.

Session shape belongs here too: timing regularity, the absence of the incidental navigation real users generate, form-fill cadence, and whether the action sequence matches how the product is actually used.

Timezone, locale, and claimed geography

Geolocation agreement proves nothing by itself, because geo-matching is a purchased feature. Internal consistency still does.

The IP resolves to a place. The browser reports a timezone offset. The client sends Accept-Language, and the OS exposes a locale. Those should agree, and for real users they usually do. A login from a residential address in Ohio, from a browser set eight hours away, with a language preference matching neither, is a contradiction even though every individual value is plausible. Comparing against the account's own history is stronger still: the question is not whether this location is suspicious in the abstract but whether this account has ever behaved this way. Our IP location checker resolves the geographic half of that comparison.

The stack, not the signal

No single layer is sufficient, and that is the real answer to why IP blocklists fail. They are not wrong. They are one layer carrying a decision that needs several.

A clean address, a TLS fingerprint that does not match the claimed browser, a device fingerprint already seen on thirty other accounts, a locale contradicting the geography, and a velocity pattern visible only once you stop counting per IP: any one alone is a maybe. Together they are a decision. Practitioners working this problem land in the same place, correlating network signals with device evidence, behavioral evidence, and downstream business context such as chargebacks, account changes, and cross-merchant takeover patterns, because IP reputation and page-load signals alone no longer give reliable visibility into this attack class.

Frequently Asked Questions

Yes, but not by looking at the IP address. The exit address is a real consumer ISP address and will usually pass reputation and ASN checks. Detection comes from the layers the proxy does not change: the TLS handshake fingerprint, the HTTP/2 connection preface, the browser and device fingerprint, the consistency between claimed geography and reported timezone and locale, and velocity measured per account or per device rather than per IP.

Trueguard Basic is free.

Start identifying visitors and signals right away, for free

Sign up for free

No credit card required.

trueguard-logo© 2026 Trueguardinfo@trueguard.io