Stockholm, Sweden / Platform engineering

ROMAN
USOV.

Senior Platform Engineer at Pierce

From the physical rack to AI operations.
I follow the problem through every layer, then leave behind a system other people can run.

Three investigations. One way of thinking.
The whole column.Compute → Network → Cloud → AI
AT PIERCE / 24MX · XLMOTO · Sledstore15+ years in technology10+ years owning production infrastructure
01 / Pierce · Network source of truth

A network you
can query.

Every useful infrastructure question starts the same way: what is actually connected to what?

What the model holds
A network you can queryConceptual view: topology, device configuration, addressing and physical paths brought together as one NetBox model of VLANs, cables and ports, with switch configuration backups behind it. Not a production topology.TopologyDevice configAddressingPhysical pathsMany inputsNetBoxONE SOURCE OF TRUTHVLANsCablesPortsTraceable relationshipsSwitch configuration backups
Conceptual model, not the production topology.

I traced the estate and modelled it in NetBox. VLAN by VLAN. Cable by cable. Port by port.

Tracing beats transcribing: a model is only worth having if it is complete. I made it something any engineer can query, with configuration backups behind the switch fleet.

The estate can be reasoned about
before it is touched.

Paths traced, changes reviewed, improvements proposed from the model rather than from memory.

NetBoxNetwork modellingConfiguration backups
02 / Pierce · Edge caching rules

Same URL.
Same 200.
Different body.

I own the caching rules at the edge — what gets cached, for whom, and for how long, across older parts of the stack carried through a platform migration. What makes them worth real attention is that a status code tells you nothing about what was actually served: the same URL can come back full or light depending on the requesting agent and whether the cache is warm. So I read response bodies, per agent, warm and cold.

Recognised crawler200 OK
Navigation. Products. Copy. Links. A fully rendered page.
Other user agents / cold cache200 OK
SHELLA response body without the rendered page
Cache and origin rules can combine to serve a full body to a recognised list of agents and a lighter one to everything else. Same status code either way — which is why the body, warm and cold, is what I check.

Conceptual response comparison. No production page or response body is reproduced.

The method

Read what the server actually returned. Compare the bodies, not just the status codes or keyword dashboards.

The finding

Cache behaviour is infrastructure, and it decides what the outside world actually sees. Caching is where an infrastructure setting turns into a commercial outcome, so it is worth one person owning it deliberately — especially while a stack is mid-migration and two generations of it are live at once.

03 / Pierce · Linux fleet

Security updates,
by procedure.

The same pattern applied to packages and patching: Red Hat Satellite and Pulp as the source of truth for what exists on which machine, and a defined procedure for getting security updates onto the fleet.

A patch round is a repeatable operation with an auditable result: content versioned deliberately, promoted in order, and answerable per machine. The platform upgrade and the readiness for the next operating systems came out of the same discipline.

Red Hat Satellite + Pulp / Security update procedure
  1. SyncAdvisories and repositories pulled in
  2. StageContent versioned, not taken live on trust
  3. PromoteReleased to the fleet in order
  4. AuditPatch state answerable per machine
RepeatableThe same steps every round
DeliberateContent chosen, not inherited
AnswerableEvidence per machine
Satellite 6.13 → 6.19Readied for RHEL 10Readied for Rocky 10
Pierce / Further outcomes

Work that stays working.

The test is what remains after the change: known state, repeatable operations and fewer things that depend on memory.

DB

Large monitoring dataset recovered and kept healthy by database tuning.

Pierce
PKI

Certificate lifecycle automated end to end: issue, store, renew, deploy.

Pierce

Sole infrastructure on-call through Black Week, seven days straight.

Pierce · 2024 & 2025
1 estate

Migrated from Akamai to Cloudflare, including edge routing and Terraform reconciliation.

Pierce · 2026
Scope / Physical to operational

The whole column.

Depth at one layer helps. Following a problem across the boundaries is where I do my best work.

08

AI operations

AT PIERCE

Internal LLM service deployed company-wide.

INDEPENDENT INFRASTRUCTURE

Multi-agent operations, persistent memory and scheduled health rounds.

07

Cloud

AT PIERCE

AWS organization built from the root account up. Entra ID SCIM, dual-active IPSec design, Azure and Cloudflare Workers.

06

Platform

AT PIERCE

Docker and Harbor. Harvester 1.7 / RKE2 proof of concept on nested vSphere. API-gateway evaluation and ADR: Kong against KrakenD.

INDEPENDENT INFRASTRUCTURE

Talos Kubernetes, Cilium and Kube-OVN.

05

Observability

AT PIERCE

Zabbix 7.0 LTS with automated proxy deployment. ELK major upgrade, Grafana, Kafka consumer lag via Burrow and node_exporter.

INDEPENDENT INFRASTRUCTURE

VictoriaMetrics.

04

Edge & mail

AT PIERCE

Akamai to Cloudflare migration, Workers routing and Terraform state reconciled with production. SPF, DKIM and DMARC across all domains, with SPF flattening.

03

Network

AT PIERCE

NetBox modelling, switch configuration backups, firewall changes by API, VLANs and IPSec/VPN.

INDEPENDENT INFRASTRUCTURE

VyOS routing.

02

Storage

AT PIERCE

NetApp NFS/CIFS. Blade-controller latency traced to NFS nconnect tuning.

INDEPENDENT INFRASTRUCTURE

LINSTOR / DRBD and ZFS.

01

Compute

AT PIERCE

VMware / ESXi on Cisco UCS, Linux, Red Hat Satellite and Pulp.

INDEPENDENT INFRASTRUCTURE

Proxmox.

Build it so the next person
does not need you in the room.

The secrets and certificate platform follows that rule: deployments pull secrets at deploy time; certificates renew and deploy themselves. Rotation becomes a repeatable process, rather than a favour asked of the person who built it.

Independent / Built in my own time

After the day job.

Different constraints. The same engineering standards.

01 / Own infrastructure

An operations layer
for a team of one.

Talos Kubernetes on Proxmox. Cilium and Kube-OVN networking, LINSTOR/DRBD and ZFS storage, VyOS routing.

The constraint is one person and no rota. The operations workflow is designed to turn collected state into a decision about what needs human attention.

Scheduled agent workflow
  1. Collect state across the layers
  2. Compare with the previous report
  3. Triage what changed
  4. Decide whether a person is needed
  5. Page with the evidence attached
From an item to a listing
OWN PRODUCT / LIVE

Uselist

A Telegram-first inventory and listing assistant. Photograph an item, get recognition and a ready-to-post listing, publish to marketplaces.

In production

Real users, Stripe billing and eBay integration. Python, Telegram Bot API, PostHog and Grafana.

A living body of knowledge
OWN PROJECT / NON-COMMERCIAL

Batrachos

A Ukrainian biology and ecology portal. Migrated from Drupal to Django and Wagtail through a rehearsed cutover, with no acceptable downtime window.

Visit Batrachos
Gear, movement and upkeep
OWN PRODUCT / WORKING ALPHA

Layrova

A kit and upkeep companion for cyclists and runners. Photo-based gear cataloguing, weather-aware clothing and component wear from ride data.

Current stage

A working alpha. An ongoing product, not a finished release.

Roman Usov / Stockholm, Sweden

Good systems start
with good questions.

English · Ukrainian
Swedish (improving)