All personas
Portrait of Marcus

Persona 04 · Emerging

Marcus, the platform engineer

Pragmatic SRE/platform engineer at an AWS/Kubernetes scale-up, evaluates cautiously, secondary persona today.

Archetype
SRE / DevOps / platform engineer
Pronouns
he/him
Typical plan
Team → Enterprise
“Every vendor says 'AI finds anomalies'. Fine. Show me the false-positive rate and the IAM policy, in that order.”
— Marcus

The pain

  1. 1

    Threshold tuning is a permanent tax

    Every new service means new alarms, and static thresholds are wrong the moment traffic shifts.

  2. 2

    Silos with AI sprinkled on top

    Every vendor's AI sees only its own data. None of them read the code or the runbooks.

  3. 3

    45 minutes of archaeology per incident

    What changed, which deploy, which config, every single time, with a platform team of two.

The aha moment

Reading an investigation and catching himself nodding: it checked what changed first, distrusted the easy answer, and wrote it up like a postmortem he would sign.

How we make them anticipate it

  • Speak his discipline back to him: silence is a feature, evidence over vibes, every action on the record.
  • Publish real investigations verbatim, wrong turns included; he converts on how it thinks, not what it claims.
  • Invite the audit: read-only connect, diff it against your own alarms, mark the homework yourself.
Status: secondary, emerging persona. The product deliberately has no enterprise motion yet, and several tools Marcus considers table stakes (PagerDuty, Grafana, Splunk, GitLab, Terraform-aware infrastructure linking) are still on the roadmap. Marcus is where the product expands next, not where it wins today. He matters now because he is the veto vote when Laura's company grows into his.

Snapshot

FactDetail
Age30-45
RoleSRE, DevOps, or platform engineer
CompanyScale-up of 30-200 engineers
Team sizePlatform team of 1-3, serving everyone else
ExperienceTen-plus years: one datacenter, one failed Kubernetes migration, three observability vendors
OwnsThe on-call rotation's quality: alert precision, runbooks, postmortems
TellReads the IAM policy before clicking connect; enters as the person vetting a tool someone else brought in

A day in his life

Starts with the overnight alert review: what fired, what should not have, what stayed silent that should not have. The morning is tickets from product teams (a new service to onboard, a quota to raise, "why is my pod pending"); the afternoon is the actual project work: Terraform refactors, cost optimization, the migration that has been 80% done for a quarter. Interruptions are the job, which is why he automates with the persistence of someone who has been paged for the same thing twice. He maintains the internal wiki nobody else updates and the dashboards everybody else screenshots.

Stack and tools

ToolThe job it does
Go / Python / BashGo for tooling, Python for glue, Bash for everything he swears is temporary
AWSIn depth: Lambda, ECS, EKS, RDS, SQS, the works
TerraformEverything as code, with opinions about state management
KubernetesEKS clusters, a service mesh he regrets
Helm / Argo CDThe deploy machinery
DockerStill the unit of everything
PrometheusMetrics he trusts because he runs them
GrafanaThe dashboards everybody else screenshots
CloudWatchThe floor of observability
Datadog / HoneycombThe paid tier of visibility
OpenTelemetryThe instrumentation standard he is slowly migrating everything to
SplunkInherited, resented, still running
PagerDutyThe pager, escalation policies tuned by hand
VaultSecrets, properly
GitHub / GitLabHeavy CI, runbooks in Markdown next to the code where they belong
SlackThe incident channel and the paging bot

The homelab mirrors production more than he admits: k3s on mini PCs, Proxmox, self-hosted everything, monitored better than some companies' production.

Goals and motivations

  • Make pages rare, actionable, and evenly distributed. The rotation's health is his professional pride.
  • One live topology across AWS, Kubernetes, and the edge platforms, with change history attached to every node.
  • Detection that learns per-resource baselines instead of static thresholds, with cadence and cost he can control.
  • Keep his existing stack: alarms and saved queries that feed the new system rather than being replaced by it.
  • Strict, auditable autonomy boundaries. He does not object to automation; he objects to automation he cannot explain in a postmortem.

Where he hangs out

WhereWhat it is to them
Redditr/devops, r/sre, r/kubernetes, r/aws, and r/homelab for love, not work
Hacker NewsA weakness for outage postmortems and "how we run infrastructure" posts; he comments there more than anywhere else
Community SlacksKubernetes Slack, CNCF channels, HangOps, a platform-engineering community, two vendor Slacks joined for support, stayed for the gossip
Mastodon or BlueskyMore than X these days; the sysadmin diaspora moved and he moved with it
GitHub issuesFiles reproductions polite enough to get fixed
NewslettersSRE Weekly, DevOps Weekly, Last Week in AWS: Corey Quinn for the jokes, stays for the judgment
PodcastsScreaming in the Cloud, the Kubernetes Podcast, Ship It

He reads incident postmortems recreationally, plus the engineering blogs of Cloudflare, fly.io, and Honeycomb, and Google's SRE books, which he quotes selectively and distrusts institutionally. In person: KubeCon or re:Invent when the company pays, SREcon or Monitorama when he pays attention, local meetups where he occasionally speaks about something that broke.

How to reach him

What worksWhy it lands
Deep technical contentHow detection actually works, what the agent queries, a worked incident with the evidence trail shown; he will read 3,000 words if they are honest
Transparency artifactsThe exact IAM policy and CloudFormation template, ServiceAccount permissions (get/list/watch and nothing else), retention answers, published limitations
Trusted voicesWar stories and postmortem culture; a Corey Quinn read or an SRE Weekly link, never a display ad
A free evaluation pathTwo weeks on a low-stakes account without talking to anyone
Conference hallwaysKubeCon, SREcon, Monitorama, staffed by engineers who can answer the third follow-up question

What fails: "AI-powered" as the headline (he has a bingo card), hiding the security model behind a sales call (if he cannot find the permissions before signup, he is gone), claiming to replace his stack (a tool that ingests his ten years of tuned alarms beats a tool that dismisses them), and any pressure tactic.

What he resonates with

  • Least privilege as a worldview: read-only by default, every write gated and on the record, rollback only where explicitly enabled. "The line is reversible versus not" is a sentence he would write himself.
  • Blameless postmortem culture, runbooks as code, boring technology, Terraform plans reviewed like PRs.
  • Evidence over vibes: he trusts a system that shows the hypotheses it tested and what it ruled out, and distrusts any system that only shows conclusions.
  • Tools that respect prior investment: CloudWatch alarms picked up and investigated on arrival, Datadog and Honeycomb saved queries becoming scheduled checks.
  • Honest engineering marketing: published limitations, real numbers, changelogs that admit regressions.

What turns him off: magic, mandatory dashboards, vendors that call him a "DevOps" as a noun, per-host pricing, AI features that cannot cite their sources.

Hobbies and personality

  • The homelab is the signature hobby: self-hosted services, a rack in the garage, DNS ad-blocking for the household, monitoring dashboards for the aquarium. He automates his home with the same rigor as production and twice the joy.
  • 3D printing, hiking or long-distance cycling, homebrewing or serious coffee, board games with a rules-lawyer streak he keeps mostly in check.
  • Mentors juniors patiently and maintains one small open-source tool with three hundred stars and his exacting standards.
  • Personality: skeptical, thorough, dryly funny, allergic to hype but not to change. Slow to trust, permanent once converted, and the loudest internal advocate a vendor can have when the product earns it.

Journey through the product

He enters as the evaluator, not the discoverer, and the make-or-break moment happens on the docs site before he even creates an account: reading the connect model and the IAM policy.

The evaluation itself is a shadow run: the feed diffed against his own alarms, false positives counted, quiet resources checked for staying quiet. He ends up at home in the power surfaces most users never open: OTel ingest with his own telemetry tokens, scoped API keys, kubectl through the agent, investigation-depth tuning per tier.

Feature affinities

  • The topology graph with cross-provider stitching: a DNS record in one provider linking to a service in another is the map he has been maintaining by hand in a wiki.
  • Change intelligence on every sync: change records with impact levels, and a thread that watches the migrations he currently babysits personally.
  • Per-resource baselines with the tier cadence dial (critical resources checked every 10 minutes, low-tier daily): detection cost he can reason about.
  • Advisories as a free audit: configuration gaps, missing observability, reliability risks per resource.
  • Enterprise controls: bring-your-own LLM key and gateway, volume credits, data-plane options.

Objections and friction

  • Missing table stakes for his world: PagerDuty, Grafana, Splunk, GitLab, and Terraform-state-aware linking are not live yet. The repo-to-infrastructure linking reads platform config files, not Terraform, and he will notice within the first hour.
  • Verdict trust: he will run the product in shadow mode against his own alarms and believe the comparison, not the marketing.
  • Data governance: where telemetry lands, how long it is retained, who can query it, and what leaves his boundary. He asks these before the security team does.
  • Unattended writes: not for a long time, and never globally. Per-account and per-action opt-ins are the only acceptable shape.

Money and plan behavior

  • Does not hold the budget but holds the veto. His written evaluation is what turns a Team plan into an Enterprise conversation.
  • Cares about cost mechanics more than cost level: predictable flat pricing with visible usage and a capped overage passes; anything that scales with hosts or ingested GB triggers the Datadog reflex.
  • The Enterprise levers that move him: volume credits, bring-your-own gateway, deployment and data-residency options, and a support channel with engineers in it.

What success looks like

  • The two-week shadow evaluation ends with a false-positive count that rounds to zero and at least one real regression caught with a correct diagnosis before his own alarms fired.
  • The rotation's mean time to understanding drops because incidents arrive with the first hour of archaeology already done and shown.
  • He retires a page of hand-tuned alert rules without losing coverage, writes it up for the internal wiki, and becomes the product's most credible advocate in the company.

In their own words

  • "Read-only role, scoped token, or no deal."
  • "I don't need it to be smart. I need it to be right, and to show its work."
  • "Ingest my alarms before you tell me to delete them."
  • "If it pages my team wrongly at 3am once, we're done. If it saves them once, I'll defend it in the budget meeting myself."