Janitor
Post

what scale actually taught me about infra

talos

this is more of an infra diary than an announcement.

i’ve been thinking about how to concepts like “running software at scale” outside of the context of how to design systems with pretty architectural diagrams.

prior to my kubernetes experience, i had no use case for it. now, i’m gradually learning the same principles that kubernetes abides by, with other components in my stack.

rke2 was where i learned to breathe

my first production ready kubernetes setup was rke2, along with rancher.

i was truly in unchartered territory. i felt like i was constructing a space shuttle with various blog posts. i had no clue what a storage class was or what ingress was. the ship (or space shuttle) was rapidly sinking, and so was my mental health.

it took a lot of trial and error and clicking around rancher to understand what was going on, and how to ensure that my services were exposed to the outside world.

with rancher’s gui tooling, i learned how to be less ignorant, and stopped having to take wild stabs in the dark to figure out what was going on.

even though rke2 has its issues, it is an amazing learning resource, and opens up the world of production ready kubernetes.

k3s made it lighter

later, i set up k3s.

again, i felt less lost. it had the same potential as rke2 to open up the world of kubernetes for me, but with less overhead.

i use ansible to configure machines. it helped automate provisioning, but it only defines the initial configuration. it doesn’t constrain the state of the system. over time, the configuration of the machine and the provisioning script differ. the script defines a desired state. the machine defines a real state. the difference between the two is revealed during an incident.

i’ve fired k3s and ansible from the job. they’re both great.

if you’re simply handling a low volume of requests, a system made up of a couple machines cleverly interconnected by ansible, may be good enough. however, if you find yourself handling a larger volume of requests and your system is still made up of ansible machines, you may start questioning your life choices.

the nature of handling a larger volume of requests often involves tuning parameters to tweak how the system manages concurrency. other concerns may include distributing load across a larger number of machines, keeping machines in a consistent state, and ensuring that multiple processes running on a single machine remain coordinated.

talos clicked

eventually you’ll begin to develop an appreciation for systems that remain boring at scale. talos is one such operating system. its purpose is to serve kubernetes. it claims to be the first operating system of its kind. talos’ creators promised to remove the peripheral purpose from machines running kubernetes, and talos delivers. the absence of an interactive shell may feel constraining at first, but it may be reassuring to someone who has managed a cluster of kubernetes nodes at scale.

configuration in talos is declarative. node configurations are serializable. you can check them in, update them, and recreate them. if something is changed outside of the talos control plane, it is obvious where the change originated. if a node dies, it is replaced.

talos and kubernetes fit well together. kubernetes focuses on desired state. talos moves this to the lowest level of the system, the kernel.

this allows kubernetes to be as easy to deploy as other talos based workloads.

talos is used for our kubernetes control plane. talos is also used to build our node workloads.

bare metal is the direction

a fleet of 50 extremely fast, and extremely powerful, amd epyc worker nodes, allows us to achieve low latency and high throughput for our workloads. this gives us better economics than the hyperscalers.

we have a 50gbps connection to the internet. egressing data on this connection costs us. talos allows us to provide a low latency, high throughput service at a lower cost than the hyperscalers.

a cloud provider is great for low latency, high throughput services. the cost of egressing data from the cloud quickly changes cloud computing from being a great service, to an uncompetitive, and frustrating experience.

caching, carefully

we have used cloudflare’s edge cache and origin server to provide low latency and high throughput to our customers. there were always rules to never serve stale assets to our customers. we learned we cannot always trust a cdn. this was most evident when cloudflare served pages to users that included the sensitive state of other users.

cdns can provide great service for truly public assets. cdns should not be trusted to provide secure, user specific assets.

our cache is farther from the services, and aggressively caches private data. stay cautious when dealing with personal info. we assume there are cache invalidations for every write. there is no true ttl.

the dragonfly part

we use bullmq heavily, and it processes the majority of our background jobs. it does message processing and a ton of other stuff, like user notifications, cleanup, andretries. it also processes moderation actions. additionally, it handles asynchronous tasks and spikes during deploys. bullmq is great, but without proper diligence, users will notice a build up of messages.

plain redis does not use multiple cores, and becomes a bottleneck. dragonfly is an upgrade from plain redis and completely removes the bottleneck.

dragonfly also is easier to manage than redis sentinel or cluster modes.

load shedding instead of dying

additionally, we have learned you can’t always block bad actors.

there are tons of obvious, easy to block abuse cases. however, there are even valid user cases where a bad actor is sending thousands of requests per second.

the focus shifts from “try to keep everything online” to “prioritize keeping users online, especially the important ones, and let some dribs and drabs of traffic die off.”

one approach to traffic management is to implement real load-shedding mechanisms at the edge. turnstile, along with cloudflare, can be used to implement rate-limiting. expire resources to limit the time spent waiting for requests by cheaper endpoints. protect more expensive endpoints using a throttling mechanism. use circuit breakers to fail a dependent service if one of its backends takes too long to respond. implement back-pressure mechanisms to allow priority traffic to pass. limit the length of requests, especially background requests.

failing a request is often better than requiring users to wait for a response. it is better to respond with a status code 429, indicating that the user has been rate limited, than for the server to take 30 seconds to respond with a 500 status code.

the goal is to keep the website usable.