SELECT_ENGAGEMENTS · SENIOR_LEVEL_ONLY

Hard data problems.
Solved properly.

If a source you depend on keeps blocking you — or the data you need lives behind an enterprise bot wall — that's the specific problem I solve. Reliable structured-data feeds that survive vendor updates, plus the pipeline and backend around them. Built from production experience, not tutorials.

Evidence_Of_Work

→

Day-job scale: 10–15M raw items/day at peak across 120+ crawlers in 23 markets

→

Single targets up to 1.7M–3.5M live listings, kept current

→

Access that survives enterprise anti-bot — Cloudflare, DataDome, PerimeterX

→

Mobile-only sources unlocked from APK to production feed

→

Zero-data-loss pipelines on PostgreSQL + AWS

→

Two live products shipped solo — architecture to deploy

What I Take On.

RELIABLE DATA ACCESS

When a source you depend on keeps blocking you.

You need data from a site that doesn't want to give it up — a competitor's catalog, a marketplace, a pricing surface behind an enterprise bot wall. Off-the-shelf scrapers work for a week, then the vendor ships an update and the feed goes dark. The expensive part isn't the scraper; it's everything downstream that quietly starves when it stops.

I build access as an architecture, not a trick. The crawler is engineered to survive enterprise anti-bot (Cloudflare, DataDome, PerimeterX) by matching what a trusted session actually looks like across the whole stack — so a vendor update breaks one layer, not the feed.

In my current role I operate this at production scale: 10–15M raw items/day at peak across 120+ concurrent crawlers in 23 European markets, with single targets up to 1.7M–3.5M live listings kept current. The full method is in the case studies.

How the anti-bot architecture works →

Deliverables

  • →A feed that survives anti-bot updates instead of breaking on them
  • →Coverage of a target that off-the-shelf scrapers can't reach
  • →Sustainable request cadence — access that lasts, not a burst that gets banned
  • →Mobile-only sources unlocked (APK → API → crawler)
  • →A clear read on what's actually extractable before you commit budget

STRUCTURED, TRUSTWORTHY DATA

Raw crawl output isn't data yet.

Pulling pages is the easy 20%. The other 80% is turning messy, duplicated, multi-format HTML into records your systems can trust — deduplicated, normalized, currency-aligned, and delivered on a schedule you can plan around.

I design the pipeline end to end: change-tracking stores in PostgreSQL, canonical deduplication, campaign state machines with crash recovery so a failed run never loses or double-counts data, raw HTML archived for replay, and a delivery step that outputs whatever your downstream needs.

The result is a feed you can build on — new markets configurable by a non-engineer, zero-data-loss recovery, and predictable delivery.

How the pipeline scales →

Deliverables

  • →A clean, deduplicated dataset in the schema your systems expect
  • →Weekly or on-demand delivery you can plan around
  • →Zero-data-loss recovery — a crashed run resumes, nothing double-counts
  • →New markets and sources added without pulling in an engineer
  • →AWS infrastructure that scales with volume (ECS, SQS, S3, RDS)

BACKEND THAT RUNS IN PRODUCTION

The systems around the data.

Data work rarely stops at the crawler. You need the APIs, billing, auth, dashboards, and deployment around it — built to run unattended, not to demo once.

Architecture-first Node.js/TypeScript: NestJS APIs, Next.js full-stack, PostgreSQL, Stripe, Docker, CI/CD, and Cloudflare/AWS deploys. Typed, tested at the boundary, deployable in minutes.

Two live products built solo, end to end: Al Bayrouni (NestJS + Next.js + PostgreSQL + Stripe) and the APEX crawl platform. Both in production.

How the auth system is built →

Deliverables

  • →Production APIs — typed, tested, documented
  • →Auth, billing (Stripe), and dashboards wired end to end
  • →A deploy pipeline your team can ship through safely
  • →Infrastructure you can hand off without it falling over

How It Works.

01

Scoped Problem

Every engagement starts with the specific problem — not a service package. What data you need, from where, at what frequency, into what system.

02

Fixed Scope & Outcome

Deliverables, timeline, and the exact data format you receive are defined before work starts. No open-ended retainers, no scope creep by default.

03

Solo by Design

No team, no handoffs, no junior doing the real work. The person on the call is the person who architects and ships. One owner for the whole problem.

04

Selective

Engagements are sized to the problem, not billed by the hour to tire-kickers. Rates follow scope on a short call. Bring the problem first.

OPEN_SOURCE

The components are public.

The detection and transport tooling behind this work is open source: whichwaf, ja3lab, impersonate, tokenbank, app2api, driftwatch — pinned on GitHub. The engineering you'd be hiring is inspectable before you hire it.

START_HERE

Got a data problem?

Describe the problem — what you need, from where, what the target looks like, and what you want on the other end. No commitment, no pitch. If it's a fit, a scoping call follows.

INITIALIZE_CONTACT

RAHMOUNIDEV · SELECT_ENGAGEMENTS

© 2026