Redesigning Tenwrite's Observability Stack
Migrating Tenwrite's telemetry from different vendors to our own self-hosted stack.
Tenwrite has multiple components deployed in a distributed manner. The main web app is served from Cloudflare Workers while the database and the backend are hosted on infrastructure we own. Then, there is the Tenwrite Google Workspaces Add-on that runs entirely on the browser and has a Google App Script and a sidebar! Conventionally, we achieved visibility by using Hetrix for monitoring the API state, Sentry for web app errors and Google Cloud Logging for proxy and backend logs.
Three different services, each an entry in the expenses book, helpful during troubleshooting but lacked something. It was evident when we had to debug a sheet automation issue, and we spent hours sifting through logs and cross-referencing events. We didn’t have what we need under a single pane of glass. For my own apps and demos, I have a dedicated observability stack setup, based on open source software. I figured, why not bring Tenwrite’s telemetry into it? Granted it would be a bit of work, but it would save us some expenses and let us customize exactly what we want to see.
The Setup
The stack is the regular LGTM stack with OpenTelemetry - only Mimir is replaced with Victoria Metrics - I guess LGTVM stack? An OpenTelemetry collector runs as a gateway and gets the telemetry data to their respective backend based on the signal type. Logs go to Loki, traces to Tempo and metrics to Victoria Metrics.
To get the data flowing from Tenwrite’s network to the observability network, I was initially considering transfer over HTTPS with authentication but realized soon that it would leak credentials - anyone inspecting network traffic from the web app could easily get the bearer token! After a small discussion with Rupam, it turns out we had already solved this problem before - I just didn’t know it, LOL. We decided the safest approach is to open a private tunnel between the two networks - WireGuard came in handy for this. So basically, we set up an encrypted connection between the Tenwrite and the Observability network and all telemetry data and probes that need to communicate across the landscape goes through this.
Instrumentation and Collection
With the tunnel in place, all that was left to do was adding instrumentation to the code-bases and setting up an OpenTelemetry collector as an agent in the Tenwrite network. It has one goal - collect and send all telemetry data to the gateway collector via the private tunnel. We used the OpenTelemetry libraries to instrument the backend code. As for the web app and add-on, we send the telemetry data to a collection endpoint the backend exposes. It is a non-blocking endpoint running separately from the main backend app, and it simply translates the incoming data into OpenTelemetry language and forwards them to the agent for collection.
A common pitfall when instrumenting distributed systems is linking telemetry data to specific events or causes. For example, a user logs in, connects a site, does a sync then performs an export. All of them emit telemetry individually - but while troubleshooting or understanding a user journey bottlenecks, each provide a part of the larger picture. Unless we link them together, we retain the age-old problem of manually sifting through tons of data. This is called correlation in observability terms - grouping distinct telemetry signals under a single root or event. We took correlation very seriously. As a result, we first decided on a standard naming convention for signals and what fields to keep as labels - cardinality was a concern so we had to be careful.
Time to Battle-Test
A few weeks after the self-hosted stack was online, a user reported an issue. I was excited - not because I get to find a flaw in the system to fix but because I wanted to see if all the effort really pays off. For the first time in a long while, it felt good to not have to write things down and switch tabs. I felt nostalgic as it reminded me of the old TCS days. I had the end-to-end system visible right in front of me. Correlations paid off and with the user’s email, I was able to track down exactly where the system tripped for him. Turns out, the system behaved as it should, and it correctly rejected an export that was too big. Our support SLA is 24 hours but this time, we were able to respond to the user much sooner than usual after the first acknowledgement.
It’s been a while since we moved to our self-hosted stack. Over the months, the system has helped us see and fix subtle bugs we didn’t know existed before users could come across them. It also monitors network ingress and helps us track malicious attempts on the network. You won’t believe how much your network is bombarded with these attempts unless you see it. Which makes perimeter defense a critical part of self-hosting. More on that in a separate blog. We have blocked and reported multiple IPs, thanks to this.