Fractional Data Engineer

How We Approach Data Engineering (Data Is Water, We're the Plumbers)

Our approach in one line: data is water, we're the plumbers, and there's a big difference between pipes that technically carry water and pipes built to last. Here's how that shows up in what we build, how we test, and how we communicate.

AO
Aline Oliveira
Published August 12, 2026 · 4 min read

philosophy · testing · communication · data engineering · fractional data engineer

Quick answer

We think of data as water and data engineering as plumbing. Getting water from a source to a tap is easy to do badly and hard to do well, and data is no different. You can build pipes that technically work and leak everywhere, or pipes built to last with the right pressure and no waste. We aim for the second kind: low-cost by design so the system stays cheap as its value compounds over time, tested for actual accuracy instead of "nobody's complained yet," and built alongside clients through weekly, low-friction communication. Here's what that looks like in practice.

Anyone can move water. Not everyone builds it to last

Data as water, data engineering as plumbing, is a metaphor we come back to a lot, because it's accurate in a way most data metaphors aren't. Getting data from point A to point B isn't the hard part. You can hack together something that moves it, the same way you can run a garden hose through a window and call it plumbing. It'll carry water. It'll also leak, and it'll be the wrong solution the moment you need more than a trickle.

Elegant plumbing is sized for the load, routed sensibly, and built so a problem in one section doesn't take down the whole house. The parallel here is a pipeline that's well-structured and efficient, built by someone who thought about what happens when the volume triples, not just whether it runs today.

We aim for low cost from the start, specifically, not as a vague value statement. Data volume grows, the number of people querying it grows, the number of things built on top of it grows, and whatever inefficiency is baked in at the start gets multiplied by all of it. So we think about tool choice (do you need the expensive option, or does the cheaper one do the job), scheduling (does this need to run every five minutes, or is hourly fine), and algorithm efficiency (a query that takes 90 seconds instead of 9 because nobody indexed the right column adds up fast once it runs dozens of times a day). None of it is exciting to talk about. All of it decides whether a system gets cheaper to run relative to its value over time, or quietly becomes a cost problem nobody budgeted for.

We test for accuracy, not for silence

Our default is the highest accuracy we can reasonably deliver, not "ship it and see what clients notice." Testing up front costs less than finding out a number's been wrong for two months because someone finally cross-checked it.

That said, accuracy and speed are sometimes in tension, and we don't pretend otherwise. If a client needs something quickly, we can release at lower accuracy on purpose, but they know that going in. It's a decision made together, not a corner cut quietly: we flag the trade-off, ship the faster version, and come back to raise the accuracy once the urgent need is handled.

From the team

Curious how this looks applied to your data?

We can walk through what low-cost, well-tested infrastructure would actually look like in your setup.

See how fractional works →

Weekly, in writing, before we talk

We start with discovery, usually a session or two, to understand the goals and the gaps before proposing anything. After that we move to a weekly cadence, and the format matters as much as the frequency.

The day before each weekly session, an email goes out with three sections: What We Accomplished, Items That Need Your Attention, and What's Next. Clients read it ahead of time instead of hearing it cold in a meeting. The live conversation focuses on the first two, mostly working through anything in "needs your attention" that requires an actual decision. What's Next often shifts based on that conversation, priorities get communicated, decisions get made, and the plan adjusts instead of running on autopilot.

That's active listening applied structurally, not just as a meeting skill. Everything that needs a client's input is named explicitly, ahead of time, in writing, so nobody's guessing what's happening or getting surprised by a decision they weren't part of.

Why any of this matters

When data is easier for employees to get to, without digging through five tools or waiting on one person, they spend more of their time doing their actual job. When systems are cheap to run and accurate, the business can trust decisions made from them, whether that's client retention, growth analysis, or reliably delivering what was promised on time.

Low-cost plumbing, tested for accuracy, built with the client in the loop every week. Not because it's a nice philosophy to have, but because it's what makes data infrastructure worth the money spent on it.

Getting this built

This is the same approach behind every engagement we run, whether it's a greenfield build or stepping into a stack that already exists. A fractional data engineer applies it end to end and typically costs a fraction of a full-time hire.

Start with the roadmap below to see what this looks like for your specific setup.

FAQ

What does 'data as water, data engineering as plumbing' actually mean?

Water gets from a source to where someone needs it through pipes, and how well those pipes are built determines the pressure, the leaks, and the long-term maintenance. Data is the same. It gets from a source system to the person who needs it through pipelines, and how well those are built determines whether the numbers show up fast, correctly, and without someone having to babysit them.

Why do you prioritize low-cost systems from the start?

Because data volume and usage only grow, so whatever you build gets run more often and on more data over time. A tool choice, a schedule, or an unoptimized query that's cheap at low volume gets expensive fast at scale. Choosing efficient options early costs a little more thought upfront and saves real money every month after.

How do you handle accuracy versus speed when a client needs something fast?

We default to the highest accuracy we can deliver, but when a client needs something quickly, we'll ship at lower accuracy on purpose, as long as they know that's the trade-off. It's a conversation, not a surprise. They decide with full information, and we come back to raise the accuracy once the urgent need is met.

What does a typical communication cadence look like?

A discovery session or two at the start to understand goals and gaps, then weekly after that. An email goes out the day before each weekly session with three sections: What We Accomplished, Items That Need Your Attention, and What's Next. The live conversation focuses on the first two, especially anything that needs a client decision, and What's Next adjusts based on what gets decided.

1 of 3 client spots remaining

Want to see this approach applied to your stack?

Answer a few questions and get a personalized data roadmap in under 5 minutes.

Get your 5-minute data roadmap →