Two-Phase Cloudflare Bypass on a $0.72/month AWS EC2 Instance
A lightweight, self-bootstrapping AWS system that monitors 11 premium matcha products for stock availability. Uses a two-phase fetch strategy to bypass Cloudflare: Scrapling StealthyFetcher launches headless Chromium every 25 minutes to solve Turnstile and extract cookies (~200 MB peak RAM), then curl_cffi polls every 60 seconds with Chrome TLS fingerprint impersonation (~58 MB steady-state RAM). Runs as a systemd service on a free-tier t2.micro instance provisioned entirely via Terraform with embedded user_data — no SSH needed.
Premium matcha products from Marukyu Koyamaen frequently go out of stock with unpredictable restocking schedules. The site is protected by Cloudflare Turnstile, making simple HTTP scraping impossible. Manual checking is tedious and misses brief availability windows.
Designed a split scraping strategy that avoids keeping Chromium resident in memory. The solve phase launches headless Chromium via Scrapling every 25 minutes to obtain cf_clearance cookies, then immediately closes it. The poll phase uses curl_cffi with Chrome TLS impersonation and cached cookies for lightweight 60-second checks. The entire stack is provisioned by Terraform and self-bootstraps via cloud-init user_data — no manual SSH required.
Explore the main capabilities and functionality of this project
Chromium solve phase (25 min) + curl_cffi poll phase (60s) keeps steady-state RAM at ~58 MB - Fits within t2.micro's 1 GB RAM
Scrapling StealthyFetcher solves Turnstile challenges and extracts cookies for lightweight reuse
Instant alerts with product name, price, status change, and direct link on stock changes
Single terraform apply provisions VPC, EC2, Lambda scheduler, and fully configures the instance via cloud-init
Key challenges faced during development and how they were solved
Technologies and tools used to build this project
Splitting browser-heavy tasks from lightweight polling dramatically reduces cloud resource requirements
curl_cffi with Chrome TLS fingerprint impersonation is sufficient for Cloudflare-protected sites when paired with valid cookies
Cloud-init user_data with gzip compression is a powerful pattern for self-bootstrapping EC2 instances
Choosing Telegram over Discord was pragmatic — Discord blocks AWS IP addresses (HTTP 403)
Proactive cookie refresh (25 min vs 30 min expiry) prevents silent monitoring failures
I'd love to discuss this project in detail and share insights about the development process.