Fixed price · 3 working days · remote

Production Infrastructure Risk Review — 3 days

See what could fail first—before your next release or enterprise customer.

An independent, three-day review of your production infrastructure. You get a concrete risk register and a prioritized action plan—with no cloud, tool, or license sales.

  • 15+ yearsproduction operations
  • 100+ serversunder direct responsibility
  • 50+ platformsoperated
  • German GmbHfor DE/EU
Review areas

Six questions the review checks systematically

This is a review checklist, not a claim about your environment. Every observation is tied to traceable evidence.

Designed for product companies with 11–200 people, software and product agencies, AI consultancies preparing for launch, and teams without a dedicated senior infrastructure owner.

01 · Recovery

Has restore actually been tested?

Backup jobs can succeed while recovery time and data integrity remain unverified.

02 · Dependencies

Where is a single point of failure left?

The review covers individual services, nodes, providers, and knowledge concentrated in one person.

03 · Signal

Does monitoring warn early enough?

Dashboards and metrics are checked for whether they prompt a clear action before users are affected.

04 · Delivery

Do deployment and rollback work?

Manual steps, approvals, dependencies, and the usable path back are made visible.

05 · Access

Are access, secrets, and edge controlled?

The review checks growing permissions, secret flows, attack surfaces, and operating rules at the edge.

06 · AI Operations

Is the LLM/GPU workload operable?

For AI systems, the review looks at capacity, isolation, observability, fallbacks, and operational routines.

Process

Three days from system context to prioritized decisions

The review begins once scope is confirmed and the agreed read-only access is available.

Kickoff and boundaries

60 minutes with the technical owner: goals, architecture, critical paths, known constraints, and access model.

Read-only review

Architecture, deployment, access, backups, monitoring, edge/security, failure modes, and runbooks are reviewed.

Report and readout

You receive the risk register and 30/60/90-day plan, followed by a 60-minute CTO or leadership readout.

Deliverables

Working material your team can act on—not a wall of slides

01 · Risk register

Prioritized findings

Severity, potential impact, evidence, and a recommended next action for each confirmed risk.

02 · Plan

30/60/90 days

An order of work that separates near-term risk reduction from structural improvement.

03 · Readout

Decision-ready context

60 minutes with the CTO or leadership to cover evidence, trade-offs, ownership, and open decisions.

Example risk-register entry

How a risk-register entry is structured

The example below shows the format and level of detail used in the report.

Open sample entry EX-01

Example data: Names and identifying details are not included. Findings in a commissioned review are based on the reviewed environment and traceable evidence.

Risk ID / Area
EX-01 · Backup & recovery
Severity / Priority
High · P1 / 0–30 days
Risk
Recoverability of the primary database has not been demonstrated by a current restore test.
Example evidence
Automated backup jobs report success. In this sample, the runbook contains no date, result, or measurements from a completed restore.
Potential impact
Recovery time and achievable data state remain uncertain after a database failure.
Owner
Platform owner
Next action
Restore a representative backup in an isolated environment, record RTO/RPO results, and schedule a recurring test.
Scope

What the fixed price includes—and what it does not

Included

  • 60-minute kickoff with the technical owner
  • Read-only review of the agreed production infrastructure
  • No more than two systems or environments
  • Up to eight hours reviewing materials and system context
  • Risk register with severity, impact, evidence, and action
  • Prioritized 30/60/90-day plan
  • 60-minute readout for the CTO or leadership

Not included

  • Penetration testing or active security testing
  • Compliance certification or formal ISO audit
  • Implementation or remediation of findings
  • 24/7 incident response or on-call coverage
  • A promise of “zero risk”

€2,490 excl. VAT covers this fixed scope. Anything beyond it is scoped separately in advance. An optional hardening/launch sprint typically runs 5–10 days and can be agreed at fixed scope or €95/hour; the review does not commit you to follow-on work.

Infrastructure engagement · 2023–2024

From fragile container deployments to safer product releases

A European product company needed a reliable way to ship several applications without turning every release into an infrastructure risk.

Starting point and delivery

Applications ran in Docker containers through a cumbersome CI/CD process. There was no standardized deployment platform or usable rollback path, while infrastructure costs were increasing.

Sergey designed and built a 20-node Kubernetes cluster, migrated Symfony and Go applications, rewrote the CI/CD pipelines, and added rollback and monitoring.

  • Kubernetes
  • Redis Cluster
  • PostgreSQL / Patroni
  • ClickHouse
  • RabbitMQ
  • Kafka
  • Symfony
  • Go
Private AI client engagement

Generate customer communications without sending confidential data outside

A client commissioned an internal content-generation system that used private customer context while keeping data, models, and outputs entirely on the client's own infrastructure.

Data and model path
  1. Approved contextOnly customer and communication data permitted for the specific draft was retrieved internally.
  2. Private retrieval layerRAG supplied relevant context from a self-hosted index while access and sources remained traceable.
  3. Self-hosted inferenceAn internally operated model generated the draft through vLLM. Prompt, context, and output remained inside the client's infrastructure.
  4. Review before deliveryPolicy checks and a human reviewer verified content, tone, and facts. The system did not send autonomously.
Reviewed by Sergey Pikalev

Production practice across infrastructure, data, and AI workloads

For more than 15 years, I have designed, built, and operated production systems—from edge infrastructure through Kubernetes and databases to private LLM inference.

I work directly with your technical owner during the review. There is no cloud-provider or license sale, and no handoff to a junior delivery team.

  • Kubernetes
  • Terraform
  • PostgreSQL
  • Edge / DDoS
  • Observability
  • vLLM
FAQ

Practical questions before the scope call

What access do you need?

Read-only access is enough. If direct access is not possible or preferred, a shared-screen session with your technical owner can be agreed.

Can we sign an NDA first?

Yes. An NDA can be agreed before confidential architecture or operations information is shared.

Do you work white-label?

Yes. For software, product, and AI agencies, the review can be agreed as a white-label engagement.

Is this a penetration test or ISO audit?

No. This is a technical and operational risk review. It does not replace a penetration test or compliance certification.

Can you fix the findings?

Implementation is not part of the review. A separate hardening/launch sprint with its own scope can be agreed afterwards.

Where do you work?

The review is delivered remotely and contracted through VML Development GmbH.

Next step

Clarify scope in 20 minutes

Tell me briefly about the system and timing. I usually reply within one business day with a proposed time for the scope call.

  • No budget field and no long questionnaire
  • Discuss read-only access or screen share
  • Confirm scope and start window before engagement
Do not include passwords, secret keys, or personal data.

By submitting, you send these details to VML Development GmbH so the enquiry can be handled. See the Privacy Policy (German) for details.