Staff Platform Engineer · SRE · DevOps — Vancouver, WA
Matt Friesen
Staff Platform Engineer
- Vancouver, WA
- [email protected]
- GitHub
Staff platform engineer and DevOps leader with extensive experience in cloud infrastructure, automation, CI/CD, and security compliance. Proven ability to lead teams, modernize infrastructure, optimize system performance, and implement best-in-class DevOps practices. Adept at building scalable solutions and empowering teams through automation and strategic process improvements.
- $100K/moAWS spend cut
- 99.999%uptime · PCI-DSS
- SOC 2Type II & GDPR
- 40%faster deploys
- AWS · GCPKubernetes · Terraform
Highlights
Skills
Experience
Staff Platform Engineer · TopstepI lead platform engineering across AWS and GCP (Kubernetes and GitOps, reliability and SLOs, observability, AI tooling, and incident response) while working at the full-stack level, delivering vertical slices from frontend to infrastructure.
Cut AWS spend by over 20%20%+ AWS savings ($100K/mo)
I analyzed resource usage and cost across AWS, GCP, Datadog, and GitHub at the organizational level, then right-sized and eliminated waste. This cut AWS spend by over 20%, about $100K per month, before counting any Savings Plan commitments. I also feed that usage data into capacity planning decisions.
Modernized infrastructure from EC2 to KubernetesEC2 → Kubernetes + GitOps
I led the move from EC2-hosted services to Kubernetes, with ArgoCD handling GitOps deployments and GitHub Actions replacing CodeDeploy. I also set the organization's CI/CD strategy and standards so every team's pipelines are consistent, secure, and reliable instead of each team reinventing its own.
Drove AI adoption with LiteLLM and org-wide standardsLiteLLM + shared AI standards
I rolled out LiteLLM as the organization's shared gateway to LLM providers, and I established standardized skills repositories and AI usage standards across engineering, so teams build on common, governed tooling instead of one-off setups.
Delivered full-stack vertical slices, frontend to infrastructureFull-stack, not one lane
Rather than staying in a single lane as an infrastructure, DevOps, backend, or frontend engineer, I work at the full-stack level and deliver vertical slices end to end: frontend, backend, CI, and the infrastructure underneath. Seeing the whole path is what let me improve local development and streamline onboarding, so new engineers get productive sooner and day-to-day work has less friction.
Built the platform engineering practice around SLOsSLO/SLI framework adopted
I defined the SLO/SLI framework that product teams adopted and now use to make tradeoff decisions, along with the incident response protocols, on-call standards, and operational runbooks behind it. I partner with product engineering early in design so reliability is built in from the start instead of retrofitted after launch.
Owned observability on Datadog and CloudWatchLow-noise, actionable alerts
I instrumented distributed tracing and closed the visibility gaps that slowed down diagnosis of production issues. I rebuilt alerting so it catches real problems without paging people for noise, and I keep tuning the signal-to-noise ratio as the platform changes.
Designed AWS and GCP infrastructure in TerraformMulti-account IaC
I design and evolve Topstep's cloud infrastructure across AWS and GCP, balancing reliability, security, scalability, and cost efficiency. I set the infrastructure-as-code strategy so every AWS account is provisioned the same way with Terraform, which makes environments consistent, repeatable, and reviewable, and I inform capacity planning for infrastructure and platform services.
Led incident response and blameless post-mortemsOutages → systemic fixes
I lead incident response and run blameless post-mortems that turn outages into systemic improvements. I established and keep refining incident classification, escalation paths, and communication protocols, and I set documentation standards for infrastructure architecture, operational procedures, and platform decision records.
- Cut AWS spend by over 20% (~$100K/month), excluding Savings Plan commitments, by analyzing usage and cost across AWS, GCP, Datadog, and GitHub, right-sizing resources, and eliminating waste.
- Led the migration from EC2-hosted services to Kubernetes, with ArgoCD GitOps deployments and GitHub Actions replacing CodeDeploy; set org-wide CI/CD standards so every team's pipelines are consistent, secure, and reliable.
- Drove AI adoption by rolling out LiteLLM as the shared LLM gateway and establishing standardized skills repositories and AI usage standards across engineering.
- Delivered full-stack vertical slices across frontend, backend, CI, and infrastructure; improved local development and streamlined onboarding so new engineers get productive sooner.
- Defined the SLO/SLI framework product teams use for tradeoff decisions, along with incident response protocols, on-call standards, and operational runbooks.
- Instrumented distributed tracing and rebuilt alerting on Datadog and CloudWatch to catch real problems without paging on noise.
- Designed AWS and GCP infrastructure with a multi-account Terraform strategy for consistent, repeatable, reviewable environments; inform capacity planning for platform services.
- Lead incident response and blameless post-mortems; set incident classification, escalation, communication, and documentation standards.
Staff Site Reliability Engineer · Streem (Frontdoor)Fast-paced startup acquired by Frontdoor in 2020. Zero-install web software and native SDKs for WebRTC video calls, with embedded AR and AI features. Technical lead for infrastructure, the CI/CD pipeline, and backend services.
Led SOC 2 Type II and GDPR complianceSOC 2 Type II & GDPR
Developed and implemented the policies and procedures to achieve SOC 2 Type II and GDPR compliance, and built security best practices for AWS services into the platform for strong data protection. Automated most of the ongoing compliance work with Strike Graph, significantly reducing manual effort.
Migrated CI/CD twice: CircleCI → GitHub Actions → GitLab CI40% faster deploys
Led two CI/CD migrations, first from CircleCI to GitHub Actions and then to GitLab CI. Both succeeded, improved developer experience, and cut deployment time by 40%.
Restructured and modernized AWS infrastructureAll infrastructure in Terraform
Rebuilt the AWS infrastructure as code with Terraform, and built a robust delivery pipeline on GitHub Actions, AWS CodeDeploy, and CodePipeline.
Designed region redundancy and centralized authenticationMulti-region platform
Designed and implemented region redundancy and centralized authentication for the platform.
Overhauled secrets managementShort-lived, scoped secrets
Designed a shorter-lived secrets lifecycle with more tightly scoped permissions and secrets injected at runtime.
Built backend microservices in KotlinDesign → production
Designed, wrote, tested, and deployed microservices in Kotlin.
- Led SOC 2 Type II and GDPR compliance: wrote the policies and procedures, built AWS security best practices into the platform, and automated most ongoing compliance work with Strike Graph.
- Led two CI/CD migrations (CircleCI → GitHub Actions → GitLab CI), improving developer experience and cutting deployment time by 40%.
- Rebuilt AWS infrastructure as code in Terraform, with delivery pipelines on GitHub Actions, CodeDeploy, and CodePipeline.
- Designed region redundancy and centralized authentication; redesigned secrets management around short-lived, tightly scoped secrets injected at runtime.
- Designed, built, and deployed backend microservices in Kotlin.
Senior DevOps Engineer · InComm Digital SolutionsIDS is the backend behind most digital gift card and reloadable credit card transactions: thousands of transactions per second in a PCI-DSS environment that must stay at 99.999% uptime.
Enabled zero-downtime deployments0-downtime deploys
Implemented HAProxy backed by Consul for dynamic routing to microservices (blue/green, percentage-based, and geo). This enabled zero-downtime deployments and fully unlocked CI/CD.
Drove monolith-to-microservices modernizationMonolith → microservices
Led the architectural strategy and decision-making behind a successful product modernization, turning a monolithic system into a scalable, efficient microservices architecture.
Ran security compliance for PCI-DSS and SOC systemsPCI-DSS · 99.999% uptime
Established and maintained the security compliance protocols for PCI-DSS and SOC-certified systems.
Rolled out HashiCorp VaultSecrets, auth & internal CA
Implemented Vault for secrets management, authentication, and a certificate authority for the microservices.
Rolled out HashiCorp Consul for service discoveryLower cross-region latency
Implemented Consul for service discovery. This reduced network latency across geographic locations, added fault tolerance, and gave services a consistent configuration store.
Codified infrastructure with AnsibleNo more snowflake servers
Moved new and existing infrastructure to code with Ansible, eliminating one-off configuration issues and long-lived legacy servers.
Created a CD readiness programOpt-in continuous deployment
Built the program dev teams went through to opt in to the fully automated continuous deployment pipeline.
- Implemented HAProxy with Consul-driven dynamic routing (blue/green, percentage-based, geo), enabling zero-downtime deployments and fully unlocking CI/CD.
- Led the architecture strategy for modernizing a monolithic system into microservices.
- Established and maintained security compliance for PCI-DSS and SOC-certified systems.
- Rolled out HashiCorp Vault (secrets, authentication, internal CA) and Consul (service discovery, lower cross-region latency).
- Codified infrastructure with Ansible, eliminating one-off configuration and long-lived legacy servers.
Cloud Operations Engineer · Viewpoint Construction SoftwareDelivered Viewpoint's construction software as a fully managed platform for thousands of customers in virtualized environments, administering and automating more than 2,000 servers across multiple datacenters and clouds.
Built a billing database that fixed billing accuracy+$500K revenue, $120K saved/yr
Built a customer management database that resolved billing discrepancies in a pay-for-use model, generating an additional $500,000 in annual revenue. The same database made Microsoft SPLA licensing reports more accurate, saving $120,000 per year.
Rolled out Datadog across 4 datacenter providersProactive operations
Implemented Datadog monitoring across AWS, GCP, and two private datacenters. This made operations proactive, improved visibility, and reduced downtime.
Lead engineer on product releases, including Viewpoint Enterprise CloudProduct launches
Worked with dev, QA, and product teams to develop the architecture, delivery methods, and deployment pipelines for the release.
Automated Microsoft SPLA license reporting$120K saved per year
Used the CMDB to report SPLA licensing more accurately, saving the department $120,000 per year.
Built an automation framework for deployments and maintenanceReused across teams
Developed an automation framework of PowerShell deployment packages, reusable scripts, and scheduled jobs covering group policies, API calls, SQL tasks, migrations, software upgrades, and maintenance, which system admins and support reused in their workflows. Also implemented infrastructure as code with PowerShell DSC.
Built cloud billing dashboards and alertsExec cost visibility
Built dashboards, budgets, and alerts with CloudWatch, Stackdriver, Google Data Studio, and BigQuery, giving executives insight into usage for resource allocation and cost optimization.
Administered 400+ SQL Server instances400+ databases
Handled backups, troubleshooting, maintenance jobs, and performance tuning.
Received 2 SPOT awards2× SPOT award
Awarded by upper management for going beyond the job description and for excellence in the workplace. Also provided tier-3 support for clients when needed.
- Built a customer management database that fixed pay-for-use billing discrepancies, adding $500K in annual revenue and saving $120K/year on Microsoft SPLA licensing.
- Rolled out Datadog across AWS, GCP, and two private datacenters, making operations proactive and reducing downtime.
- Built a PowerShell automation framework for deployments, upgrades, and maintenance that sysadmins and support reused; implemented infrastructure as code with PowerShell DSC.
- Built cloud billing dashboards and alerts (CloudWatch, Stackdriver, BigQuery) that gave executives cost visibility.
Technical Consultant · Synergy Business SolutionsMicrosoft Dynamics ERP implementations, upgrades, and customizations for a Microsoft Presidential Award–winning reseller.
Migrated company to Azure and Office 365On-prem → cloud
Moved on-premises servers, Active Directory, and email to Azure and Office 365.
Advised clients on server and system architecture50%+ billable
Specced new servers for clients and advised on system architecture. More than half of all work was billable, client-facing work.
Database administration and reportingSQL Server 2000–2014
Administered Microsoft SQL Server (2000 through 2014) and wrote reports.
Ran internal IT and SharePointAD, DNS, firewall, SharePoint
Managed Active Directory, web servers, DNS, DHCP, and the firewall. Administered SharePoint 2007, 2010, 2013, and SharePoint Online, and set up virtual environments for testing software implementations.
- Migrated on-premises servers, Active Directory, and email to Azure and Office 365.
- Specced servers and advised clients on system architecture; more than half of all work was billable.
Professional Soccer Player (Center Midfield) · Royal Liège & Kitsap PumasPlayed for Royal Liège in Belgium (2008–09) and the Kitsap Pumas, later Sounders 2 (2009–2012).
2011 USL champion and 2012 USL Best XIAll-time leading scorer, Kitsap
Won the 2011 USL championship, became Kitsap's all-time leading goal scorer, and was named to the 2012 USL Best XI.
- 2011 USL champion, Kitsap's all-time leading goal scorer, and 2012 USL Best XI.
Education
B.A. Computer Science · Whitworth UniversityGPA 3.8
- Academic Excellence Scholarship for the best GPA in the major (2007–2008)
- Education Award for the Act Six and service-learning programs (May 2008)
Certifications
- AWS Certified AI Practitioner (AIF-C01) · Amazon Web Services
- Administering Microsoft SQL Server 2012 Databases (Exam 70-462) · Microsoft