Field Signal · Curated Reading
The Engineering
Signal.
High-signal field notes, architecture teardowns, and engineering writing from 50 publications. A focused reading queue for becoming a stronger end-to-end engineer. Updated 20 Aug.
Claude Opus 5 on GitLab: Reasoning built for the hard tasks
A mistake on a routine task can cost you a minute. A mistake on a large refactor or a debugging trail spanning months of commit history can cost far more, as it compounds silently over hundreds of exchanges. By the time you catch it, every step built on top of it needs unwinding, too. That's the dif...
Normalize security logs to Google SecOps UDM with Observability Pipelines
Learn how Observability Pipelines normalizes your telemetry to Google SecOps UDM, enabling both consistent investigations across sources and precise upstream control over your SIEM ingest.
ModelExpress: Distributing Model Artifacts at the Speed of Light
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse, moving...
Turn And Face The Strange
We’re Fly.io, a public cloud platform that is both our favorite way to put an app on the Internet and our favorite way to safely let a frontier agent coding harness cook. This is a post about our company, the future, and Sprites, which are computers for agents that you can check out right now. This ...
Postgres backups under the hood
With backups being such a vital part of keeping your data safe, how do they actually work?
A new allowlists design for Grafana Cloud IP addresses: What you need to know
If your network restricts inbound or outbound traffic, you likely maintain an allowlist of Grafana Cloud IP addresses so your systems and Grafana Cloud can talk to each other. Today we're introducing a new allowlists design : a single, structured API that replaces the collection of per-product lists...
Terraform Stacks, explained
Terraform Stacks simplify provisioning and managing resources at scale, reducing the time and overhead of managing infrastructure.
Terraform introduces workspaces and Stacks restore, and more
HCP Terraform and Terraform Enterprise boost resiliency, governance, and scalability, helping platform teams manage infrastructure and deliver faster at scale.
The Pulse: New trend - concern about massive increase in code review load
Top of mind for engineering leaders: what to do about the growing code review load, and how devs are starting to review code less thoroughly than before? Many questions, but few proven solutions.
Debugging Ray Tracing Applications Using NVIDIA OptiX Toolkit
NVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways...
Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a...
A Beginner’s Guide to Clocks, Causality, and Ordering in Distributed Systems
Why does something as simple as reading the time become a hard problem for distributed systems?
How to monitor your Supabase projects: connect Grafana Cloud in one click
As AI agents accelerate software development and spin up applications at scale, visibility into what's happening behind the scenes, including query performance and database health, has never been more important. Gaining that level of insight requires observability that can keep pace. Supabase and Gr...
What's new in Postgres 19
`VACUUM` reclaims dead-tuple space but doesn't shrink a table. PostgreSQL 19 Beta 2 adds an in-core online rewrite with REPACK (CONCURRENTLY), and much more.
Drag, drop, done: a visual composer for Redpanda Connect
A first look at the new visual Pipeline Builder for Redpanda Connect (preview), plus a roundup of recent CDC and connector updates.
Launching Health in ChatGPT
Health in ChatGPT now lets eligible U.S. users securely connect medical records and Apple Health to get more personalized insights and better understand their health.
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
Open the source for the full engineering note.
Exercises in benchmarking and evals, part 7: DeepSWE, Senior SWE-Bench, napkin math, and winter tires
This is part of a series of exercises on benchmarking, evals, and experimental design ( 1 , 2 , 3 , 4 , 5 , 6 ) 1 . We're going to look at three questions, which are presented before the answers to give you time to think about the questions before seeing the answers. 29. A friend of mine is reviewin...
Provision Datadog on Stripe Projects
Learn how to provision Datadog on Stripe Projects, start a 14-day free trial, and manage billing through your existing Stripe account.
AI gateway best practices: Model routing, reliability, and budget controls for production agents
Learn how AI gateways help you scale your agents to consume multiple LLM services reliably, and how to monitor these systems to ensure you’re getting the best performance and cost.
Offloading I/O to Dedicated Cores: An Asymmetric io_uring Backend for Seastar and ScyllaDB
We moved low-level I/O execution off application cores to dedicated networking cores using Seastar’s new asymmetric_io_uring backend. Explore the architecture design, trade-offs, and benchmark results.
Building a serverless AI assistant at Pelago: concept to care in two weeks
Healthcare organizations face a critical scaling challenge – how to maintain deeply personalized patient interactions as member bases grow, without overwhelming care teams or compromising quality. At Pelago, a digital health company specializing in substance use disorder support, the engineering tea...
Make Long-Running NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++
A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can...
Best Practices for Building AI Agents That Work in Production
In this article, we try to explore the collective thinking into a smaller set of practices and explain the reasoning behind each one, rather than asking anyone to memorize a numbered list.
Cost attribution in Grafana Cloud: Manage spend across observability and testing workflows
Knowing what you're spending on observability is useful. Knowing which team, service, or project is driving that spend is what actually lets you act on that information. Cost attribution is a core part of how Grafana Cloud approaches cost management and optimization . It gives your organization a cl...