HTTP Automation and Web Scraping

Use requests and BeautifulSoup to automate HTTP interactions, extract data, and perform basic web reconnaissance.

Project: Scraping-as-a-Service API (with MCP future path)

A common pentest/OSINT pattern is: you write a one-off scraper, then you want teammates (or other tools) to call it. Wrapping the scraper in a small HTTP API turns it into a reusable service.

Architecture

┌─────────────┐     HTTP      ┌────────────────────┐     ┌──────────────┐
│   Client    │ ───────────── │  FastAPI service   │────▶│ Target site  │
│  (curl,     │  GET /scrape  │  - rate limiting   │     │ (your target)│
│  MCP tool, │               │  - caching         │     └──────────────┘
│  another   │               │  - proxy rotation  │
│  script)   │               │  - result shaping  │
└─────────────┘               └────────────────────┘

Minimal FastAPI wrapper

from fastapi import FastAPI, HTTPException
import requests
from bs4 import BeautifulSoup

app = FastAPI(title="Scraping-as-a-Service")

@app.get("/scrape")
def scrape(url: str, selector: str | None = None):
    try:
        resp = requests.get(url, timeout=10)
        resp.raise_for_status()
    except requests.RequestException as e:
        raise HTTPException(status_code=502, detail=str(e))

    soup = BeautifulSoup(resp.text, "html.parser")
    if selector:
        elements = soup.select(selector)
        return {"url": url, "count": len(elements), "texts": [e.get_text(strip=True) for e in elements]}
    return {"url": url, "title": soup.title.string if soup.title else None}

Run it: uvicorn scraper_api:app --reload, then test with curl "http://localhost:8000/scrape?url=https://example.com&selector=h1".

Practical additions for pentest/recon

  • Rate limiting: use slowapi or a time.sleep + token bucket so you don't hammer the target.
  • Caching: cache identical URL/selector combos with functools.lru_cache or Redis to reduce noise.
  • Proxy rotation: pass proxies={"http": proxy, "https": proxy} to requests.get.
  • Auth headers: accept an Authorization header and forward custom User-Agent/Cookie headers.
  • Error shaping: return consistent JSON errors instead of tracebacks.

Next step: expose it as an MCP tool

Once the API works, you can wrap it as a Model Context Protocol (MCP) tool so an AI assistant can call it directly:

from mcp.server.fastmcp import FastMCP

mcp = FastMCP("scraper-mcp")

@mcp.tool()
def scrape_url(url: str, selector: str = "") -> str:
    """Scrape a URL and return the text of matching elements."""
    import requests, json
    r = requests.get("http://localhost:8000/scrape", params={"url": url, "selector": selector})
    return json.dumps(r.json(), indent=2)

if __name__ == "__main__":
    mcp.run()

This keeps the scraping logic in one place and lets both humans and agents use it.

Headless browser scraping with Python

When a target page is built by JavaScript (SPAs, lazy loading, anti-scraping challenges), static requests + BeautifulSoup may not be enough. Headless browsers run a real browser engine without a visible window, so they execute JavaScript, render DOM, and can act like a real user.

When to use headless scraping

  • The page content is loaded or modified by JavaScript after the initial HTML.
  • You need to interact with the page: clicks, form submissions, scrolling, logins.
  • The site uses anti-bot checks that look for a real browser environment.
  • You want to capture network traffic, cookies, or storage state.

Main Python options

Tool Engine Best for Notes
Selenium Real browser (Chrome/Firefox/Edge) via WebDriver Maximum compatibility, legacy sites, complex interactions Heavier, needs browser + driver installed
Playwright Bundled browser engines (Chromium, Firefox, WebKit) Speed, reliability, modern features, stealth Newer API, good async support
requests-html Chromium via Pyppeteer (now largely unmaintained) Lightweight dynamic rendering Limited compared to Playwright

For new projects, Playwright is generally preferred because it handles waits, isolation, and stealth more cleanly. Selenium remains useful when you need to match a specific installed browser or use existing WebDriver tooling.

Playwright example

Install:

pip install playwright
playwright install chromium

Basic scrape of a JavaScript-rendered page:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com")
    page.wait_for_selector("h1", timeout=10_000)
    title = page.title()
    text = page.inner_text("h1")
    print(title, text)
    browser.close()

Common interactions:

page.click("button#load-more")
page.fill("input[name=username]", "admin")
page.fill("input[name=password]", "secret")
page.click("button[type=submit]")
page.wait_for_load_state("networkidle")

Selenium example

Install:

pip install selenium webdriver-manager
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

options = Options()
options.add_argument("--headless=new")
options.add_argument("--disable-blink-features=AutomationControlled")

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
driver.get("https://example.com")
print(driver.title)
print(driver.find_element(By.TAG_NAME, "h1").text)
driver.quit()

Stealth and detection avoidance

Headless browsers are detectable by default. Sites can check for:

  • navigator.webdriver === true
  • Missing plugins, languages, or permissions
  • Headless-specific user agents or window sizes
  • Behavioral signals (instant mouse movements, perfect timing)

Mitigations include:

  • Using playwright-stealth or selenium-stealth packages.
  • Setting a realistic viewport and user agent.
  • Adding random delays and realistic mouse movements.
  • Running in headed mode on a real desktop for high-stakes targets.

Never use stealth bypasses against sites you do not own or have explicit permission to test.

Performance and resource notes

  • Launch one browser and reuse contexts/pages instead of launching a new browser per request.
  • In Playwright, browser.new_context() gives you isolated cookies/storage without the cost of a new browser.
  • Close browsers explicitly to avoid leaking memory and processes.

Headless browser integration with the Scraping-as-a-Service API

You can extend the FastAPI wrapper by replacing the requests fetch with a headless browser step for JavaScript-heavy targets. For example, in Playwright:

from fastapi import FastAPI, HTTPException
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeout

app = FastAPI(title="Scraping-as-a-Service")

def fetch_with_browser(url: str, selector: str | None = None, wait_for: str | None = None):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        try:
            page = browser.new_page()
            page.goto(url, timeout=15_000)
            if wait_for:
                page.wait_for_selector(wait_for, timeout=10_000)
            elif selector:
                page.wait_for_selector(selector, timeout=10_000)
            html = page.content()
        except PlaywrightTimeout:
            raise HTTPException(status_code=504, detail="Timeout waiting for page content")
        finally:
            browser.close()

    from bs4 import BeautifulSoup
    soup = BeautifulSoup(html, "html.parser")
    if selector:
        elements = soup.select(selector)
        return {"url": url, "count": len(elements), "texts": [e.get_text(strip=True) for e in elements]}
    return {"url": url, "title": soup.title.string if soup.title else None}

@app.get("/scrape-js")
def scrape_js(url: str, selector: str | None = None, wait_for: str | None = None):
    return fetch_with_browser(url, selector, wait_for)

This gives the API a /scrape-js endpoint for dynamic pages while keeping the lightweight /scrape endpoint for static pages.

  • Only scrape sites you own or have written authorization to test.
  • Respect robots.txt, rate limits, and terms of service.
  • Headless browsers can be more aggressive than simple HTTP requests; use delays, cache, and proxy rotation to avoid causing harm.