python/

Building Python connectors that last

11 min read codestinger updated

A lot of my work is integration: code that moves data between an ERP, a manufacturing system, a warehouse system and whatever API sits in the middle, over REST, SOAP, SFTP or FTPS. None of these systems were designed with each other in mind.

The first version of a connector usually works on the developer's machine. The version that matters is the one that is still running two or more years later, at a customer site, with very little maintenance. Getting there takes three kinds of work: understanding the environment before writing code, building habits into the runtime and investing in everything that keeps the code stable over time.

Part 1: Before writing any code#

Most connector problems I've seen were decided before the first line of code: the wrong assumption about volumes, an operating system nobody asked about, a firewall rule nobody requested. So I start with questions:

  • Infrastructure: which operating systems and Python versions, which proxies, firewalls and internal certificate authorities? Can we install packages or must everything ship as a self-contained bundle?
  • Volume and latency: records per hour at a normal peak and at the worst peak of the year, message sizes and the latency the business actually needs. These decide polling or streaming, batching or single calls.
  • The slowest system: if the ERP accepts twenty requests per second, the design respects that from day one.
  • Timelines and lifetime: a smaller connector that is tested and documented on time beats a complete one held together by hope. And a connector expected to run for years in a regulated environment is designed very differently from a one-off migration script.

Part 2: Habits that survive production#

1. Every network call gets a timeout#

Many HTTP clients will wait forever by default. One hung connection and your nightly job is still "running" at lunch.

import httpx

client = httpx.Client(
    base_url="https://erp.example.com/api",
    timeout=httpx.Timeout(10.0, connect=5.0),  # 5s to connect, 10s for everything else
)

Create the client once and reuse it; it keeps connections alive and makes timeouts consistent.

2. Retry the right failures, with backoff and jitter#

Some failures are temporary: dropped connections, 429 Too Many Requests, 502/503/504. Retry those. Others are permanent (400, 401, 404) and retrying them just hammers the other side.

import random
import time

import httpx

RETRYABLE_STATUS = {429, 500, 502, 503, 504}


def request_with_retry(client, method, url, *, attempts=5, base_delay=0.5, max_delay=30.0, **kwargs):
    for attempt in range(1, attempts + 1):
        try:
            response = client.request(method, url, **kwargs)
        except httpx.TransportError:
            if attempt == attempts:
                raise
            retry_after = None
        else:
            if response.status_code not in RETRYABLE_STATUS or attempt == attempts:
                response.raise_for_status()
                return response
            header = response.headers.get("Retry-After", "")
            retry_after = float(header) if header.isdigit() else None

        # Exponential backoff with full jitter, unless the server told us how long to wait
        delay = retry_after or random.uniform(0, min(max_delay, base_delay * 2**attempt))
        time.sleep(delay)

The jitter matters more than it looks. Without it, fifty workers that failed together retry together and knock the recovering server over again.

3. Make writes idempotent#

Retries create a new problem: did the first POST succeed before the connection dropped? If you retry blindly you may create the same order twice.

Give each logical operation a stable key and let the receiver deduplicate on it. Many APIs accept an Idempotency-Key header; for systems that don't, keep a table of processed message IDs on your side.

import uuid

key = str(uuid.uuid5(uuid.NAMESPACE_URL, f"order:{order.number}:{order.revision}"))
request_with_retry(client, "POST", "/orders", json=order.payload(), headers={"Idempotency-Key": key})

Deriving the key from business data (rather than a random UUID) means a re-run of the whole job produces the same keys.

4. Stream pages and keep secrets outside#

Treat paginated APIs as a generator that yields items page by page, so memory stays flat and work starts immediately. Keep credentials in the environment or a secret manager and check them all at start-up, so a missing one fails loudly before the run begins rather than halfway through it.

5. Log like you'll debug it at 3 a.m.#

Every log line should answer which run, which record, which system. A correlation ID per run and the business key per record turn a wall of text into something you can search:

import logging
import uuid

log = logging.getLogger("connector.orders")
run_id = uuid.uuid4().hex[:8]

log.info("order synced", extra={"run_id": run_id, "order": order.number, "target": "erp"})

Pair it with a formatter that prints those extra fields (or a JSON formatter) and never log payloads that contain personal data or secrets.

6. Validate at the boundary#

The other system will eventually send you something unexpected: a missing field, a date in the wrong format, a quantity as a string. Parse and validate everything at the edge, quarantine bad records with a clear reason and keep the rest flowing. One malformed message should never stop a batch of thousands.

from dataclasses import dataclass
from datetime import date
from decimal import Decimal, InvalidOperation


class RecordError(ValueError):
    """A record that can't be processed. The message says why."""


@dataclass(frozen=True)
class Order:
    number: str
    quantity: Decimal
    due: date


def parse_order(raw):
    try:
        number = str(raw["number"]).strip()
        quantity = Decimal(str(raw["quantity"]))
        due = date.fromisoformat(raw["due"])
    except KeyError as exc:
        raise RecordError(f"missing field {exc.args[0]!r}") from None
    except (InvalidOperation, ValueError, TypeError) as exc:
        raise RecordError(f"invalid value: {exc}") from None
    if not number:
        raise RecordError("empty order number")
    if not quantity.is_finite() or quantity <= 0:
        raise RecordError(f"quantity must be positive, got {quantity}")
    return Order(number, quantity, due)


good, quarantined = [], []
for raw in records:
    try:
        good.append(parse_order(raw))
    except RecordError as exc:
        quarantined.append((raw, str(exc)))

Note the is_finite() check: Decimal("NaN") parses happily and then raises on comparison. Edge cases like that are exactly what the tests in Part 3 are for.

Part 3: Built to last#

Compatibility, including the legacy#

Integrations often run where the newest tools can't. I have shipped connectors to servers where the only interpreter was Python 2.7, because the surrounding system could not be upgraded. Respect that: avoid syntax the target doesn't support, isolate the few imports that differ between versions and test on every interpreter you promise to support. Write down what you discover about the other systems too (SOAP versions, accepted TLS versions, date formats and undocumented encodings). It is the first thing someone needs when that system is finally upgraded.

Secure by default#

A connector holds credentials to two or more business systems, which makes it an attractive target. Keep TLS verification on (trust an internal CA explicitly instead of disabling checks), use service accounts with minimal permissions, pin and scan dependencies with pip-audit and bandit in CI and never write secrets or personal data to logs or quarantine files.

Compiling with Cython#

For connectors that ship to customer servers, I often compile the core modules with Cython. The same Python source becomes a native extension module. Hot loops such as mapping, parsing and transformation usually run faster while the distributed package contains compiled binaries rather than readable source.

# setup.py
import os

from Cython.Build import cythonize
from setuptools import setup
from setuptools.command.build_py import build_py

COMPILED = ["erp_connector/mapping.py"]


class build_py_without_compiled_sources(build_py):
    """Ship the compiled extension instead of the readable .py for compiled modules."""

    def find_package_modules(self, package, package_dir):
        compiled = {os.path.normpath(path) for path in COMPILED}
        modules = super().find_package_modules(package, package_dir)
        return [(pkg, mod, path) for pkg, mod, path in modules if os.path.normpath(path) not in compiled]


setup(
    name="erp-connector",
    packages=["erp_connector"],
    ext_modules=cythonize(COMPILED, compiler_directives={"language_level": "3"}),
    cmdclass={"build_py": build_py_without_compiled_sources},
)

Declare setuptools and cython under [build-system] requires in pyproject.toml so the build also works in a clean, isolated environment. pip wheel . --no-deps then produces a platform-specific wheel. Build one per target operating system and Python version in CI (Windows runners included) and check that the wheel contains the compiled module but not its .py source. Compiled code is harder to read and modify, but it is not encryption, so keep secrets out of it all the same.

Tests for the good, the bad and the ugly#

A connector earns trust through tests that cover both positive and negative cases. The happy path proves the feature works. The negative cases prove it fails safely when the world misbehaves.

from decimal import Decimal

import pytest

from erp_connector.orders import RecordError, parse_order


def test_valid_order_is_parsed():
    assert parse_order({"number": " PO-2 ", "quantity": 2.5, "due": "2026-10-01"}).quantity == Decimal("2.5")


@pytest.mark.parametrize(
    "raw, reason",
    [
        ({"quantity": "5", "due": "2026-10-01"}, "missing field 'number'"),
        ({"number": "PO-3", "quantity": "five", "due": "2026-10-01"}, "invalid value"),
        ({"number": "PO-5", "quantity": "NaN", "due": "2026-10-01"}, "must be positive"),
        ({"number": "PO-6", "quantity": "1", "due": "01/10/2026"}, "invalid value"),
    ],
)
def test_invalid_orders_are_rejected(raw, reason):
    with pytest.raises(RecordError, match=reason):
        parse_order(raw)

Beyond unit tests, integration tests exercise the connector against a stand-in for the real system. httpx ships a mock transport that makes this easy, including the failures that are hard to reproduce on demand:

import httpx

from erp_connector.http import request_with_retry


def test_retries_until_the_server_recovers(monkeypatch):
    responses = iter([httpx.Response(503), httpx.Response(503), httpx.Response(200, json={"ok": True})])
    client = httpx.Client(base_url="https://erp.test", transport=httpx.MockTransport(lambda request: next(responses)))
    monkeypatch.setattr("time.sleep", lambda seconds: None)

    assert request_with_retry(client, "GET", "/health").json() == {"ok": True}

For bigger scenarios I run the connector end to end against a simulator that produces realistic volumes and faults. All of it runs automatically in CI on every change and on every Python version the connector supports, together with a performance test at realistic volume so a slowdown is caught before a customer sees it.

Documentation that outlives the project#

Two years from now, the person supporting the connector may not be the person who wrote it. They should find a README explaining what it connects and how data flows, a configuration reference, a runbook for deploying, upgrading, rolling back and fixing common errors and a changelog that traces every version on a customer site.

The checklist#

Before a connector goes live, I ask:

Understanding

  • Do we know the client's infrastructure, operating systems and Python versions?
  • Are the volumes, peak loads and required latency agreed and written down?
  • Does the design respect the speed limits of the slowest system?
  • Is the scope realistic for the timeline?

Runtime

  • Does every call have a timeout, with only temporary failures retried using backoff and jitter?
  • Can the job be re-run safely after a crash halfway through?
  • Is memory flat no matter how much data comes back?
  • Are all secrets external and checked at start-up?
  • Can I trace one record from source to target and does a bad record get quarantined instead of killing the run?

Longevity

  • Is it tested on every interpreter and system version we promise to support?
  • Is TLS verified, are permissions minimal and are dependencies scanned?
  • Are compiled builds produced for every target platform?
  • Do negative cases get as much testing as positive ones, with integration and performance tests on every change?
  • Could someone else deploy, support and upgrade it from the documentation alone?

When every answer is yes, the connector is ready for production. More importantly, it's ready to keep running for years with very little attention.