HTTP Automation and Web Scraping
Use requests and BeautifulSoup to automate HTTP interactions, extract data, and perform basic web reconnaissance.
Project: Scraping-as-a-Service API (with MCP future path)
A common pentest/OSINT pattern is: you write a one-off scraper, then you want teammates (or other tools) to call it. Wrapping the scraper in a small HTTP API turns it into a reusable service.
Architecture
┌─────────────┐ HTTP ┌────────────────────┐ ┌──────────────┐
│ Client │ ───────────── │ FastAPI service │────▶│ Target site │
│ (curl, │ GET /scrape │ - rate limiting │ │ (your target)│
│ MCP tool, │ │ - caching │ └──────────────┘
│ another │ │ - proxy rotation │
│ script) │ │ - result shaping │
└─────────────┘ └────────────────────┘
Minimal FastAPI wrapper
from fastapi import FastAPI, HTTPException
import requests
from bs4 import BeautifulSoup
app = FastAPI(title="Scraping-as-a-Service")
@app.get("/scrape")
def scrape(url: str, selector: str | None = None):
try:
resp = requests.get(url, timeout=10)
resp.raise_for_status()
except requests.RequestException as e:
raise HTTPException(status_code=502, detail=str(e))
soup = BeautifulSoup(resp.text, "html.parser")
if selector:
elements = soup.select(selector)
return {"url": url, "count": len(elements), "texts": [e.get_text(strip=True) for e in elements]}
return {"url": url, "title": soup.title.string if soup.title else None}
Run it: uvicorn scraper_api:app --reload, then test with curl "http://localhost:8000/scrape?url=https://example.com&selector=h1".
Practical additions for pentest/recon
- Rate limiting: use
slowapior atime.sleep+ token bucket so you don't hammer the target. - Caching: cache identical URL/selector combos with
functools.lru_cacheor Redis to reduce noise. - Proxy rotation: pass
proxies={"http": proxy, "https": proxy}torequests.get. - Auth headers: accept an
Authorizationheader and forward customUser-Agent/Cookieheaders. - Error shaping: return consistent JSON errors instead of tracebacks.
Next step: expose it as an MCP tool
Once the API works, you can wrap it as a Model Context Protocol (MCP) tool so an AI assistant can call it directly:
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("scraper-mcp")
@mcp.tool()
def scrape_url(url: str, selector: str = "") -> str:
"""Scrape a URL and return the text of matching elements."""
import requests, json
r = requests.get("http://localhost:8000/scrape", params={"url": url, "selector": selector})
return json.dumps(r.json(), indent=2)
if __name__ == "__main__":
mcp.run()
This keeps the scraping logic in one place and lets both humans and agents use it.
Headless browser scraping with Python
When a target page is built by JavaScript (SPAs, lazy loading, anti-scraping challenges), static requests + BeautifulSoup may not be enough. Headless browsers run a real browser engine without a visible window, so they execute JavaScript, render DOM, and can act like a real user.
When to use headless scraping
- The page content is loaded or modified by JavaScript after the initial HTML.
- You need to interact with the page: clicks, form submissions, scrolling, logins.
- The site uses anti-bot checks that look for a real browser environment.
- You want to capture network traffic, cookies, or storage state.
Main Python options
| Tool | Engine | Best for | Notes |
|---|---|---|---|
| Selenium | Real browser (Chrome/Firefox/Edge) via WebDriver | Maximum compatibility, legacy sites, complex interactions | Heavier, needs browser + driver installed |
| Playwright | Bundled browser engines (Chromium, Firefox, WebKit) | Speed, reliability, modern features, stealth | Newer API, good async support |
| requests-html | Chromium via Pyppeteer (now largely unmaintained) | Lightweight dynamic rendering | Limited compared to Playwright |
For new projects, Playwright is generally preferred because it handles waits, isolation, and stealth more cleanly. Selenium remains useful when you need to match a specific installed browser or use existing WebDriver tooling.
Playwright example
Install:
pip install playwright
playwright install chromium
Basic scrape of a JavaScript-rendered page:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com")
page.wait_for_selector("h1", timeout=10_000)
title = page.title()
text = page.inner_text("h1")
print(title, text)
browser.close()
Common interactions:
page.click("button#load-more")
page.fill("input[name=username]", "admin")
page.fill("input[name=password]", "secret")
page.click("button[type=submit]")
page.wait_for_load_state("networkidle")
Selenium example
Install:
pip install selenium webdriver-manager
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
options = Options()
options.add_argument("--headless=new")
options.add_argument("--disable-blink-features=AutomationControlled")
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
driver.get("https://example.com")
print(driver.title)
print(driver.find_element(By.TAG_NAME, "h1").text)
driver.quit()
Stealth and detection avoidance
Headless browsers are detectable by default. Sites can check for:
navigator.webdriver === true- Missing plugins, languages, or permissions
- Headless-specific user agents or window sizes
- Behavioral signals (instant mouse movements, perfect timing)
Mitigations include:
- Using
playwright-stealthorselenium-stealthpackages. - Setting a realistic viewport and user agent.
- Adding random delays and realistic mouse movements.
- Running in headed mode on a real desktop for high-stakes targets.
Never use stealth bypasses against sites you do not own or have explicit permission to test.
Performance and resource notes
- Launch one browser and reuse contexts/pages instead of launching a new browser per request.
- In Playwright,
browser.new_context()gives you isolated cookies/storage without the cost of a new browser. - Close browsers explicitly to avoid leaking memory and processes.
Headless browser integration with the Scraping-as-a-Service API
You can extend the FastAPI wrapper by replacing the requests fetch with a headless browser step for JavaScript-heavy targets. For example, in Playwright:
from fastapi import FastAPI, HTTPException
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeout
app = FastAPI(title="Scraping-as-a-Service")
def fetch_with_browser(url: str, selector: str | None = None, wait_for: str | None = None):
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
page = browser.new_page()
page.goto(url, timeout=15_000)
if wait_for:
page.wait_for_selector(wait_for, timeout=10_000)
elif selector:
page.wait_for_selector(selector, timeout=10_000)
html = page.content()
except PlaywrightTimeout:
raise HTTPException(status_code=504, detail="Timeout waiting for page content")
finally:
browser.close()
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
if selector:
elements = soup.select(selector)
return {"url": url, "count": len(elements), "texts": [e.get_text(strip=True) for e in elements]}
return {"url": url, "title": soup.title.string if soup.title else None}
@app.get("/scrape-js")
def scrape_js(url: str, selector: str | None = None, wait_for: str | None = None):
return fetch_with_browser(url, selector, wait_for)
This gives the API a /scrape-js endpoint for dynamic pages while keeping the lightweight /scrape endpoint for static pages.
Ethical and legal reminders
- Only scrape sites you own or have written authorization to test.
- Respect
robots.txt, rate limits, and terms of service. - Headless browsers can be more aggressive than simple HTTP requests; use delays, cache, and proxy rotation to avoid causing harm.