Book Profile
Data Management at Scale Modern Data Architecture with Data Mesh and Data Fabric, Second Edition (Final Release)
Piethein Strengholt · 2022
A pragmatic guide for large enterprises to transition from bottleneck-prone centralized data architectures to a scalable, federated model by applying principles from data mesh, data fabric, and domain-driven design.
Get the book →In an era where data is the lifeblood of business, many large organizations are shackled by outdated, monolithic data architectures like data warehouses and data lakes that stifle agility and innovation. "Data Management at Scale" confronts this challenge head-on, providing a comprehensive strategy for re-architecting your enterprise data landscape for the modern, data-driven world. Author Piethein Strengholt cuts through the hype around concepts like data mesh and data fabric, offering a pragmatic, field-tested approach that blends decentralized, domain-oriented data ownership with the essential guardrails of central governance. The book guides you through organizing your data landscape using business capabilities, treating data as a first-class product, and building a resilient architecture that supports both analytical and operational needs, ultimately empowering your organization to unlock the true value of its data and achieve scalable, sustainable growth.
What it argues
This causal model, inferred from 'Data Management at Scale', posits that adopting specific architectural and organizational design levers (Domain-Oriented Architecture, Federated Governance, Self-Serve Platforms, Data-as-a-Product) leads to improved business and architectural outcomes. This effect is mediated by fostering key psychological and behavioral states within the organization, such as clear data ownership, increased team autonomy, and greater trust in data.
Key ideas it contributes
- Domain-Oriented Architecture — The strategic design of the data landscape where data, applications, and teams are organized into decentralized, self-contained units (domains) aligned with stable business capabilities.
- Federated Data Governance — A governance model where accountability for data quality, security, and metadata is delegated to the data domains, while a central body sets global standards, policies, and provides enabling tools.
- Self-Serve Data Platform — The provision of centralized, standardized, and automated tools, services, and blueprints (e.g., landing zones) that enable domain teams to autonomously build, deploy, and manage their own data products and pipelines.
- Data-as-a-Product Mindset & Design — The practice of treating data assets as first-class products with clear ownership, SLOs, and well-defined, consumable interfaces (e.g., batch data products, APIs, events) that are designed for a positive consumer experience.
- Domain Data Ownership & Accountability — The psychological and behavioral state where domain teams feel and act responsible for the quality, security, and fitness-for-purpose of the data they produce and serve to others.
- Team Autonomy & Agility — The ability of domain teams to develop, test, and deploy data solutions independently and in parallel, with minimal dependencies on other teams or central functions, leading to faster delivery cycles.
- Data Discoverability & Trust — The state where data consumers can easily find the data they need and have confidence in its accuracy, completeness, and reliability due to clear ownership, documentation, and quality metrics.
- Data-Driven Business Value — The tangible organizational outcomes derived from leveraging data, such as increased revenue, reduced costs, improved customer satisfaction, and the creation of new data-driven products and services.