ML infrastructure + research

From prototype to production, without the rewrite.

Backprop is a two-person ML infrastructure practice. We take a problem from research and prototype through to a production pipeline - training that runs unattended, evaluation that generalizes, serving that holds up, and monitoring that catches drift.

$5,000 one week credited in full against implementation

What we do

Engineers first. Most of what breaks is infrastructure.

A model that scores well in a notebook and a model that earns its keep in production are different artifacts. We build both, and the path between them - which is usually where the work actually turns out to be.

01

Production pipelines

Training and inference that runs without a human watching it. Orchestration, reproducible builds, retries and idempotency, and a failure story you can page on. End-to-end, from raw data to a served inference.

02

Evaluation & monitoring

Offline and online evaluation harnesses, and the gap between the two measured rather than assumed. Concept and performance drift detection, alerting thresholds tied to a decision someone actually takes, and the dashboards to go with them.

03

Research & prototyping

For when the problem is still open. A first honest baseline, then the modeling that decides whether it is worth deploying at all - computer vision, NLP and recommender systems, plus fine-tuning and evaluation for LLM systems.

04

Data & platform

Lineage, freshness and leakage. Feature and label pipelines. Cloud ML platforms (Vertex AI, AzureML) and the internal services and dashboards that sit around them.

The starting point

Production readiness assessment

One week, one system, a fixed price. We review what you have, tell you what stands between it and production, and rank the work by risk. If you hire us to do that work, the fee comes off the first invoice.

$5,000FIXED

One week, start to finish. Credited in full against implementation work commissioned within 60 days - so if you go ahead, the assessment costs you nothing.

What you get

  • A written assessment: every finding ranked by risk, each with the fix and what it costs to make.
  • A prioritized remediation plan with effort estimates, in an order you can actually staff.
  • A 90-minute walkthrough with your team - engineers, not just the person who signed.
  • Any small fixes we can land inside the week, landed. You keep the code either way.

No model yet? Then the assessment is the wrong first purchase. Research and prototyping is scoped separately - tell us the problem and we will say what it would take.

01

Reproducibility

Can the model be rebuilt from data and code by someone who isn’t the author?

02

Pipeline

What runs unattended, how it fails, and how long you take to notice.

03

Evaluation

Whether the metric you optimize is the one you actually care about.

04

Serving

Latency, cost per request, model versioning and the rollback path.

05

Monitoring

Concept and performance drift, surfaced instead of left silent.

06

Data

Lineage, freshness, leakage, and the upstream change that poisons everything.

How we work

Four things we hold to.

You get both of us

The two people you meet are the two people who do the work, on every engagement, start to finish. No work goes to anyone outside the two of us, and there is no account manager in between.

Evaluation before optimization

A leaderboard is not generalization. Before anything gets tuned, we check that the metric you are optimizing is the one you care about and that your validation split isn’t quietly lying to you.

We write it down

Every engagement ends in something the next engineer can pick up: the decisions, the trade-offs, and the parts we would do differently with more time.

Working code, not slides

Deliverables run. We hand back a repository your team owns, in your stack, with tests and a way to deploy it - not a deck describing one.

Who we are

Backprop is two engineers.

One from large-scale production systems, one from competition-grade modeling. Between them, the full range: a prototype that proves the idea, and the pipeline that keeps it alive.

Muhammad Haseeb Ahmad

Production systems & ML engineering

A decade across large-scale production systems and applied ML. Built and productionized model pipelines, drift monitoring and a natural-language-to-SQL system in industry; now builds end-to-end imaging, OCR and LLM platforms in a research setting. Named inventor on a granted US patent for a clustering algorithm developed at Afiniti.

Mohibullah Kamran

Research, modeling & evaluation

Leads the research end. A Kaggle competitor who reaches the top of a leaderboard in domains he has not worked in before - third of 1,950 teams at NeurIPS 2024, on small molecule-protein binding, then medals in biomass estimation from pasture imagery and in NFL player-movement inference. The biomass work ran on a small training set, using an ensemble of DINOv3, ConvNeXt, EfficientNet and SigLIP. Formerly a fighter pilot and flight instructor, which is roughly the right instinct for production systems.

Where we have worked
  • Google
  • University of Oxford
  • The Home Depot
  • Afiniti
  • Noon Academy

These are current and former employers. They are not Backprop clients and none of them endorse Backprop. Backprop is a new practice; we show no client work, no logos and no testimonials on this site, because we have not yet earned any we are free to publish.

Get in touch

Tell us what you are trying to ship.

One email is enough to start. We will tell you honestly whether the assessment is the right first step for you, or whether it isn’t.

hello@backpropml.com

Useful to includeoptional
  1. 01The systemWhat it does and roughly how it is built today
  2. 02The gapWhat “in production” needs to mean for it
  3. 03The clockWhen it needs to be there, and what is forcing that