Grey Lathrop

Grey Lathrop

Principal SRE / Systems Architect

Contact

Email

grey@ember.st

Website

grey.ember.st

About

Principal Site Reliability Engineer and Systems Architect with 20+ years of experience designing, scaling, and operating distributed cloud-native systems on AWS. Proven record of 99.99% uptime, multi-million-dollar cost savings, and architectural leadership for engineering organizations of up to 150+. Combines deep expertise in AWS, Kubernetes, Terraform, and reliability engineering with hands-on development in Python, Rust, and Go, including production agentic AI workflows for code review, documentation, and readiness assessment. Recognized for mentoring senior engineers and translating complex technical challenges into scalable, business-aligned solutions.

Profiles

GitHub

plathrop

LinkedIn

greylathrop

Work

Torc Robotics

Senior Software Engineer torc.ai

Senior Engineer on an 8-engineer Simulation Services team, responsible for simulation infrastructure and validation tooling supporting autonomous truck testing.

Highlights
  • Managed Torc's simulation infrastructure, built on AWS primitives and processing over 50k simulation scenarios daily
  • Improved operational uptime from 90% to 99%
  • Achieved ~$30k/month (~22%) AWS cost savings through optimization
  • Architected and maintained a C++ gRPC bridge connecting the virtual driver stack to a vendor scenario engine, enabling simulated scenario validation of the autonomous driving system
  • Designed agentic AI workflows adopted team-wide: automated multi-model PR review, self-updating documentation via stacked PRs, and AI skills for production-readiness and observability review
  • Mentored engineers on infrastructure and operational best practices

MuleSoft (Salesforce)

Software Engineering Architect www.mulesoft.com

Led architecture for Production Engineering division (180+ engineers), driving operational uplift and modernization initiatives.

Highlights
  • Directed multi-year migration of MuleSoft Control Plane from self-managed Kubernetes to AWS EKS with zero downtime, improving post-migration uptime to 99.99%
  • Drove adoption of safe change practices and blue/green deployment
  • Embedded hands-on with the 3-engineer team behind a Tier-0 service depended on by all of MuleSoft, delivering a security and compliance uplift to corporate standards under deadline pressure
  • Led technical development for an organization of ~80 engineers, directly mentoring 15-20 and guiding career growth across the org
  • Led migration from New Relic to Salesforce internal monitoring platforms across production environments
  • Identified and documented multiple opportunities for multi-million dollar cost savings

Spectrum Labs

Senior Architect www.spectrumlabsai.com

Designed and deployed distributed data processing systems using Spark, Argo, and Kubernetes.

Highlights
  • Developed Terraform modules codifying best practices for EKS autoscaling and node termination
  • Reviewed major code changes for ML pipelines, improving performance and maintainability
  • Improved ingestion performance for on-prem ML classifiers by several orders of magnitude

Salesforce

Software Engineering Architect / PMTS / LMTS www.salesforce.com

Led architecture and implementation of deep observability systems for the Heroku platform.

Highlights
  • Guided hybrid deployment models emphasizing cloud coherence and shared tooling
  • Led development of system for automated provisioning of trusted AWS accounts with verifiable security properties for use across Salesforce
  • Introduced Embedded Platform Engineering model to unify platform and product teams
  • Executed lossless migration of data collection architecture handling hundreds of thousands req/s
  • Led transition from legacy CI to new CI/CD platforms with zero data loss

Krux Digital

Platform Architect / Sr. Infrastructure Engineer www.salesforce.com/products/data-cloud

Achieved consistent AWS cost controls leading to logarithmic scaling of costs, providing a key differentiator contributing to $800Mn Salesforce acquisition.

Highlights
  • Architected reorganization of DevOps and Platform teams to support scale; introduced patterns adopted org-wide
  • Designed and deployed distributed web services in Scala + Play! handling 13k QPS in production
  • Scaled Kafka clusters 50% while reducing costs 11%
  • Re-architected Java and Python systems for 10x throughput using open-source components
  • Built real-time data collection and websockets services supporting 2k data points/sec per process

Various Companies

Senior / Contract Systems Engineer

Senior and contract engineering roles at Lyft, Zicasso, Yammer, SimpleGeo, Digg, Kapor Enterprises, SquareTrade, and others.

Highlights
  • Built and automated infrastructure at scale using Puppet and FAI
  • Migrated config management systems (Chef to Puppet)
  • Introduced CI, incident response, monitoring automation, and multi-datacenter deployments

Projects

Palimpsest

Long-term semantic memory system for LLM agents: chunking, embeddings, and retrieval over persistent agent memory

Hearth

Persistent agentic runtime: identity-as-context system prompts, message routing, scheduled heartbeat, and custom context compaction

gwl-systems

Spec-first inventory and automation for self-hosted infrastructure, turning desired state into auditable, idempotent changes

gwl-agent-skills

Open collection of agent skills, including a parallel multi-model code-review pipeline with conflict resolution and feedback triage

Skills

Leadership & Strategy

  • Technical Mentorship
  • Cross-functional Alignment
  • DevOps Culture
  • Agile Methodologies
  • Communication

Cloud & Infrastructure

  • AWS (EC2, S3, EKS, Kinesis, Lambda)
  • Cloud Cost Optimization
  • Event-driven Architecture
  • Terraform
  • Kubernetes
  • gRPC
  • Simulation Infrastructure
  • CI/CD Pipelines
  • High Availability
  • Scalability
  • Virtualization

Reliability & Observability

  • Observability
  • OpenTelemetry
  • SLOs
  • Monitoring
  • Logging

AI & Agentic Systems

  • Agentic Workflow Design
  • LLM Agent Runtimes & Tooling
  • AI Skills Authoring
  • Embeddings & Semantic Search
  • AI-Assisted Development

Systems & Operations

  • Linux System Administration
  • Distributed Systems
  • System Architecture
  • Puppet
  • Automation
  • DNS

Software Development

  • Python
  • Rust
  • Go
  • Scala
  • Java
  • Common Lisp
  • Bash
  • Git

References

He doesn't skim for surface-level issues: he traces through actual execution paths and failure modes end-to-end, and never flags a problem without also bringing a concrete fix. His reviews caught subtle bugs that would only have surfaced under real production load. He's been genuinely generous with his time and judgment as a mentor.

Rohit Mishra, Software Engineer II at Torc Robotics

An exceptional software developer and architect with a rare combination of technical acumen and high-level emotional intelligence. He consistently navigates the complexities of modern engineering while keeping the human element at the forefront of every decision.

Josh Lynn, Senior Engineering Manager at Torc Robotics

A highly skilled software engineer and architect with a deep understanding of cloud orchestration workflows. He thrives in ambiguous, cross-functional environments, bringing structure and clarity to complex problems - particularly strong at articulating trade-offs and identifying risks early.

Yathartha Tuladhar, Technical Product Manager at Torc Robotics

Has the depth of knowledge that only the most seasoned systems engineers possess. Automation, packaging, Puppet, AWS, cluster management, capacity planning, etc. are mundane tasks, which is why you'll find him regularly reaching beyond normal duties for other, more interesting, challenges. A systems engineer who can program, or a programmer who knows systems engineering inside and out.

Joe Stump, Principal Engineer at Google

A true, trusted startup infrastructure guy. His great amount of experience leads him to be a good mentor. He's technically articulate, reliable in a production incident, and reasonable but passionate in decision-making. He takes ownership of what he builds and has a great ability to balance doing things the right way but within practical reason.

Ben Standefer, Founder & CEO

One of the best infrastructure engineers I've ever worked with, a fine coder as well, and someone who looks at a problem with all angles and has examined every argument before he presents a solution. I'm definitely a better engineer because I worked with him.

Jeff Pierce, Independent Technical Consultant