Overview
In this episode of The Dashboard Effect, Brick Thompson and Caleb Oaks address a gap in how most organizations think about data lake investment: the build is not the end of the work, it is the beginning of an ongoing operational commitment. The conversation covers what active data lake management actually involves, why it requires specialized skills that are different from traditional database administration, and what happens when organizations treat their data lake as infrastructure that maintains itself.
For any organization that has built or is planning to build a data lake and wants to understand what the sustained investment looks like beyond the initial implementation, this episode provides an honest and practical account of what is required. See how Blue Margin’s Managed Data Service provides the ongoing expertise and active management that keeps a data lake performing reliably as the business and its data needs evolve.
What This Episode Covers
Active Management Is Essential (0:37 – 1:07)
A data lake is a living system, not a completed project. Beyond the initial build, the platform must continuously adapt to evolving data sources, API updates, schema changes from vendors, and the pipeline failures that are an inevitable feature of any production data environment. Organizations that staff and plan for the build but not for the ongoing management find themselves with infrastructure that degrades in reliability and performance over time without a clear explanation of why.
Performance Optimization (2:09 – 3:10)
As data volumes grow, ETL processes that performed acceptably at smaller scale begin to slow down in ways that affect the freshness and reliability of reporting. Implementing delta loading, pulling only new or changed records rather than reprocessing entire datasets, is one of the most important practices for keeping pipeline runtimes efficient as the platform matures. That optimization is not a one-time configuration. It requires ongoing attention as data volumes and ingestion patterns change.
Cost Management (3:57 – 4:40)
Storing unnecessary or obsolete data creates rising cloud costs that accumulate quietly until they become significant. Anomalies like abnormally large text strings or unexpected data spikes can inflate storage costs rapidly if not monitored and addressed. Maintaining an archiving strategy that moves historical data to lower-cost storage tiers and actively managing what stays in the primary environment is as important as the technical architecture that was designed at build time.
Maintenance and Troubleshooting (4:44 – 5:53)
Pipeline failures are not exceptional events. They are a regular feature of production data environments, driven by vendors updating schemas, deprecating fields, or changing API behavior without notice. Someone needs to be responsible for receiving and acting on pipeline failure alerts promptly, because every hour a pipeline is down is an hour users are working from stale or incomplete data. That responsiveness is a staffing and process commitment, not just a technical one.
The Role of the Data Manager (6:18 – 8:31)
The person managing a data lake serves as a resident expert and gatekeeper for the platform. They maintain the data catalog, help data modelers locate and understand the data they need, and ensure that the institutional knowledge about what lives where and why does not walk out the door when personnel change. The skill set required for this role is closer to an ETL or BI developer than a traditional database administrator, which is an important distinction for organizations hiring or developing for the position.
Specialized Skill Set and Knowledge Transfer (8:42 – 10:46)
Managing a data lake requires specialized skills that are not widely distributed across the talent market. For organizations that have hired vendors to build their platforms, the hosts strongly recommend involving internal staff throughout the build process rather than receiving a finished product at the end. The knowledge transfer that happens during construction is what determines whether the organization can maintain and extend what was built, or whether it remains dependent on the vendor for every subsequent change.
Who It’s For
This episode is worth your time if you are a technology or operations leader who has invested in a data lake and is trying to understand what the ongoing staffing and operational commitment actually looks like, a data engineering team responsible for maintaining a production data environment and wanting validation and structure around the practices that keep it healthy over time, an organization that has experienced the consequences of treating a data lake as a set it and forget it infrastructure investment and is trying to understand what went wrong, or any company evaluating a data lake build and wanting an honest picture of what comes after the implementation before committing to the investment.
Why It’s Worth a Listen
The gap between what organizations are told data lake projects require and what they actually require tends to surface most painfully after the build is complete and the vendor has moved on. This episode closes that gap before it becomes a problem for organizations that are still in the planning or early implementation phase, and provides a useful diagnostic for those already experiencing the consequences of underinvesting in ongoing management.
The data manager role discussion is particularly valuable for organizations trying to staff this function. The distinction between the skills required and the traditional DBA profile is not obvious, and hiring for the wrong profile leaves a gap in the specific capabilities the platform needs to stay healthy. Understanding what to look for before the hiring process starts produces a better outcome than discovering the mismatch after the fact.
And the knowledge transfer recommendation is worth treating as a non-negotiable requirement rather than a nice-to-have in any vendor engagement. The organizations that come out of a vendor build with genuine internal capability to maintain what was built are in a fundamentally different position than those that received a finished platform without the understanding to operate it. This episode makes the case for insisting on that transfer clearly and with enough specificity to act on it.