如何在 Python 中旋转代理(requests 和 aiohttp)
轮换是指一个获取一个 IP 被封禁的爬虫和一个将负载分散到多个 IP 的爬虫之间的区别。这里有一个适用于同步请求和异步 aiohttp 的工作轮换器,带有健康检查和自动退役死去的出口。
轮换实际上解决了什么
网站通过IP进行速率限制和封禁。如果每个请求都来自同一个地址,你很快就会达到限制,单个封禁会阻止整个工作。轮换将请求分散到多个出口,因此每个IP的负载保持在阈值以下,一个失效或被封的出口只会让你重试,而不是整个运行。两种设计几乎涵盖了所有情况:轮询(按顺序循环访问池,预测性强且均匀)和随机(每个请求随机选择,简单且无状态)。
The pool itself can come from two places. You can list several s4m endpoints and ports and rotate across them, or, for quick low-stakes tests, pull live entries from the free public proxy list — filtered by country, protocol and anonymity, though those third-party proxies are use-at-your-own-risk and die often. Production rotation should lean on authenticated datacenter proxies from s4m (SOCKS5 on 1080, HTTP on 3128, USER:PASS), which stay up and are metered so scaling the pool is cheap. This builds directly on scraping with requests.
请求和aiohttp的旋转器
一个轮询的池,在失败时重试,并淘汰持续失败的出口。同步和异步版本。替换 USER:PASS。
import asyncio, itertools, random
import requests, aiohttp
from aiohttp_socks import ProxyConnector # pip install aiohttp-socks
USER, PASS, HOST = "USER", "PASS", "proxy.s4m.online"
POOL = [
f"socks5h://{USER}:{PASS}@{HOST}:1080",
f"http://{USER}:{PASS}@{HOST}:3128",
# add more endpoints/ports here to widen the pool
]
# ---------- Synchronous: requests + round-robin + retirement ----------
class Rotator:
def __init__(self, pool):
self.alive = list(pool)
self.fails = {p: 0 for p in pool}
self._cycle = itertools.cycle(self.alive)
def pick(self):
return next(self._cycle)
def retire(self, proxy):
self.fails[proxy] += 1
if self.fails[proxy] >= 3 and proxy in self.alive:
self.alive.remove(proxy)
self._cycle = itertools.cycle(self.alive or POOL)
def get(url, rot, retries=4):
for _ in range(retries):
proxy = rot.pick()
try:
r = requests.get(url, proxies={"http": proxy, "https": proxy},
timeout=(5, 20))
r.raise_for_status()
return r
except requests.RequestException:
rot.retire(proxy)
raise SystemExit("pool exhausted")
rot = Rotator(POOL)
print(get("https://api.ipify.org", rot).text)
# ---------- Async: aiohttp, one connector per proxy ----------
async def fetch(url, proxy):
connector = ProxyConnector.from_url(proxy)
async with aiohttp.ClientSession(connector=connector) as s:
async with s.get(url, timeout=aiohttp.ClientTimeout(total=25)) as r:
return await r.text()
async def main():
tasks = [fetch("https://api.ipify.org", random.choice(POOL))
for _ in range(5)]
for ip in await asyncio.gather(*tasks, return_exceptions=True):
print(ip)
asyncio.run(main())重要的设计选择
决定旋转是否有帮助或只是增加混乱的旋钮。
轮询与随机
轮询均匀分配负载,易于理解;随机是无状态和简单的。当均匀分配很重要时,选择轮询。
淘汰无效出口
计算每个代理的连续失败次数,并在几次失败后丢弃一个出口,以便池自我修复,而不是永远重试一个无效的代理。
使用前健康检查
定期通过每个代理访问 IP 回显端点,仅保留响应的代理。当池中包含不稳定的免费列表条目时,这一点至关重要。
退后,不要冲刺
轮换并不是洪水泛滥的许可证。添加每个主机的延迟和指数退避,以便您保持在限制之下,而不是在多个IP上与之抗争。
粘性会话和诚实限制
Rotation is not always what you want. Some flows — a login, a cart, a multi-step form — need the same IP for the whole session, or the site sees a user teleporting between addresses and blocks it. That is a sticky session: rotate between jobs, stay fixed within one. When a workflow needs one guaranteed-stable address, add a dedicated personal IP instead of rotating.
诚实地面对上限。s4m代理是数据中心,因此旋转它们有助于应对速率限制甚至负载,但不会掩盖数据中心IP地址,防止反机器人系统阻止整个范围。旋转对浏览器指纹或请求模式也没有任何作用。将其作为一层:与真实的头部、退避和尊重每个网站的规则配对。对于异步部分,你将需要aiohttp-socks以支持SOCKS5;使用工具验证出口,并通过API自动处理凭证。
问题,已解答
我如何使用请求轮换代理?
保持代理 URL 列表,并为每个请求选择一个 — 轮询或随机 — 作为代理参数传递。连续失败几次后退役一个出口,以便池自我修复。上面的示例显示了一个完整的旋转器。
我如何使用 aiohttp 轮换代理?
aiohttp 为 HTTP 每个请求设置代理,但对于 SOCKS5,您需要为每个代理构建一个 ProxyConnector(来自 aiohttp-socks)并使用它创建会话。使用 asyncio.gather 在池中分配任务。
我什么时候不应该轮换代理?
当工作流程需要一致的身份时——登录、购物车、多步骤表单。在会话中切换 IP 看起来可疑并会被阻止。对于这些流程,请使用粘性会话或专用 IP。
旋转的数据中心代理能避免所有封锁吗?
不。轮换分散负载并规避每个 IP 的速率限制,但数据中心范围仍可能被激进的反机器人系统整体标记,且轮换不会改变您的浏览器指纹或行为。
人们通过搜索找到此页面
- 如何在 Python 中旋转代理(requests 和 aiohttp)
- 最佳 用于抓取的代理 2026
- 用于抓取的代理 价格
- 用于抓取的代理
- 什么是 http 代理
- 什么是ISP(静态住宅)代理?
- Windows 上的 VPN 杀开关:实际有效的防火墙规则
- socks5代理 解释
- 免费代理列表是安全的吗
- HTTP 407 代理身份验证要求:如何修复它
- 免费代理列表 无需注册
- 代理认证
此页面回答的真实搜索短语 — 链接的短语打开详细覆盖它们的页面。