Data Catalogs Emerge as Strategic Requirement for Data Lakes
If the exhibitors at last week’s Strata + Hadoop World expo are any indication of what’s happening down on the street, data cataloging is evolving from a nice-to-have into a necessity for organizations looking to capitalize on big data.
Hadoop’s so-called “junk drawer” problem has been well-documented. It stems largely from the flexible schema-on-read approach, where data is structured only when it’s finally accessed from the data lake, as opposed to the traditional ETL approach of transforming data when it’s originally loaded into the data warehouse.
In short, getting data into Hadoop is easy, but finding it and getting it back out again can be hard. All sorts of vendors are now looking to address this dilemma, which touches many aspects of big data analytics, including data quality and security. Having a catalog of the data stored in Hadoop seems like a good idea, and there are a number of vendors providing that.
Alex Gorelik, CEO and founder of Waterline Data, which provides data cataloging software for Hadoop and other big data systems, says data professionals are reluctant to open Hadoop to downstream users without a better accounting of the actual data.
“The data lake looks like a flea market,” Gorelik tells Datanami. “It’s all in there somewhere, but how do you find it? It’s a problem for data scientists and data stewards because they can’t give people access until they know what’s in there.”
Gorelik says that while open source tools like Apache Atlas, which is backed by Hortonworks (NASDAQ: HDP), and Cloudera Navigator provide a good technical foundation for addressing data cataloging and master data management (MDM) challenges, they don’t go far enough to solve the problem. Waterline addresses the problem by using “tags” to track the lineage of every piece of data.
With Waterline, Hadoop users can continue ingesting data as they did before, while relying on the software to keep it somewhat organized. Apache Lucene sits under the covers to power searches, while an Amazon-like user interface and “shopping cart” process lets analysts check-out when they’ve found their data.
It’s not a license to be messy with your data, but at least it takes the burden off of users to manually track their data. “People used to have careful directories. But these days, they can’t keep track of their directories,” Gorelik says. “You have millions of files. You should organize them as well as you can. [With Waterline software] it doesn’t matter where the file is, as long as you can find it.”
Collibra is another master data management (MDM) software vendor helping customers keep track of their Hadoop-resident data using the catalog approach. The company, which recently moved its headquarters to New York City, has an eight-year history of providing data governance solution to customers in healthcare, financial services, and other industries.
“What we have is a technology platform that has the capability to keep track of processes around data, the metadata and organization and roles and responsibility for data,” Daniel Sholler, director of product marketing at Collibra, tells Datanami. “We keep track of all the technical connections of all the data because you need to know that stuff. But it turns out that stuff isn’t the interesting stuff.”
Instead, Collibra exposes a set of applications that make it relatively easy for end users to get access to data, if they are authorized to access it. That’s the “interesting” stuff that Sholler was referring to. Data access is one component of a collection of data governance solutions that Collibra is offering, and the scope of that offering will expand in the coming weeks.
Another vendor that’s plying the fruitful waters of data cataloging is Alation. The company originally designed its product to “learn” about data connections by observing how analysts interact. But just providing data cataloging wasn’t enough, the company says. So last week Alation announced that in its version 4.0 update, it will also track queries that run along with the data that’s collected.
Tracking queries and data, says Alation CTO Venky Ganti, will provide critical context that’s required for addressing the needs of data stewards and customers, including answering questions like “Where can I find data to answer my question?” “Can I trust this data?” “What are the data semantics in order to use it?” and “Who can answer my question about this data set.”
“Experts who understand certain datasets often play the stewardship role of ensuring that data consumers can make accurate and effective use of data,” Ganti says in a blog post. “More recently, data governance initiatives have started to assign formal stewardship responsibility.”
Other companies offering data cataloging functionality include Podium Data, which announced a $9.5-million Series A round just prior to the show. Zaloni also unveiled its Bedrock Data Lake Manager (DLM) product, which uses data cataloging to help manage storage more effectively. At Strata, it launched a new version of Mica, its data preparation tool, which introduces a new “shopping cart”-like experience.
That “shopping cart” metaphor was heard often on the Strata expo floor during discussions of data catalogs and big data management. You can expect to see that show up in MDM and data quality tools more often.
Informatica, the big dog of last-gen ETL tools that’s hungering for a piece of the big data pie, also updated its data lake management product, called Data Lake Management, to include more capabilities. Specifically, the product combines data cataloging, stream data capture, Hadoop job management, security, and cloud connectors in a single unified product.
The lack of a centralized data lake management point eats up analysts’ time and hurts productivity, says Amit Walia, executive vice president and chief product officer for Informatica. “Ease of use and a delightful user experience along with robust governance and metadata capabilities are critical for getting business value out of data lakes,” he says in a statement.”
According to Gartner analysts Guido De Simoni and Roxane Edjlali, enterprise metadata management, including data cataloging, has become a “required discipline.” “Failure to recognize this will lead to sustained siloed behavior and loss of business value,” they wrote earlier this year.
While data silos will inevitably be with us for a while, we don’t have to behave as if the data is trapped in a single location. As the Gartner analysts rightly point out, organizations that can get a unified view of their data will find greater business value. It’s becoming clear that data catalogs will be one way of providing that visibility.
September 30, 2020
- Machine Learning from Enea Openwave is Delivering 15% Increase in RAN Capacity
- Anyscale Announces Ray 1.0
- Perforce Releases Book About ML and AI in the Age of DevOps
- Collibra Launches New Partner Program
- GTCOM-US to provide first-of-kind APAC alternative data as part of Bloomberg’s data marketplace
- Netlist Ships Next Generation of NVMe Solid State Drives
- Kyligence and Global IT Service Provider ESS to Bring Apache Kylin and Kyligence High-Performance Analytics Solutions to Latin America
September 29, 2020
- PyTorch / XLA now generally available on Cloud TPUs
- Data Science to Accelerate Drug Discovery with Artificial Intelligence and Machine Learning, Says Frost & Sullivan
- DDN Tops the Ratings in Intersect360 User Survey for Technical and Operational Satisfaction and Future Vision for Storage
- New Denodo Platform 8.0 Accelerates Hybrid/Multicloud Integration, Automates Data Management with AI/ML, and Boosts Performance
- Intel Enters into Strategic Collaboration with Lightbits Labs
- Pepperdata Announces Query Spotlight Now Supports Apache Impala
- Oracle Helps Marketers Simplify the Management and Activation of Customer Data
- Datadobi Launches Pre-Migration Assessment Service
- Signals Analytics Awarded Wide-Ranging Patent Grant for Automatic Extraction of Information from Unstructured Data Sources
September 28, 2020
- Cohesity Announces Automated Disaster Recovery that Minimizes Application Downtime and Data Loss
- DataStax Co-Founder and CTO Jonathan Ellis to Keynote at ApacheCon 2020 on Open Source in the Cloud Era with DataStax Astra and Apache Cassandra
September 25, 2020
- PostgreSQL 13 Released: Performance Gains, Space Savings, Enhanced Security, Developer Experience
- WANdisco Announces Global Agreement with Infosys to De-Risk and Accelerate Data Lake Migration to the Cloud
Most Read Features
- How Facebook Accelerates SQL at Extreme Scale
- Big Data File Formats Demystified
- 10 Big Data Statistics That Will Blow Your Mind
- VC Ben Horowitz Dishes on Hadoop, AI, and Data Culture
- Microsoft Now Developing Its Own Hadoop
- How to Build a Better Machine Learning Pipeline
- The CDO’s Role in Leading Data-Driven Transformation
- How the Coronavirus Response Is Aided by Analytics
- The Future of Labor in an AI World
- Is Python Strangling R to Death?
- More Features…
Most Read News In Brief
- Snowflake to Make it SNOW on NYSE
- Aerospike Gives Legacy Infrastructure a Real-Time Boost
- Snowflake Pops in ‘Largest Ever’ Software IPO
- A ‘Breakout Year’ for ModelOps, Forrester Says
- Google Joins the MLOps Crusade
- Microsoft Launches Spatial Analytics, Other AI Services at Ignite
- New AI Tool Maps the Families of the Bible, A Song of Ice and Fire
- Fivetran Launches Pay-As-You-Go Option for ETL
- Air Force Expands Predictive Maintenance
- Cassandra Gets an Indexing Upgrade
- More News In Brief…
Most Read This Just In
- Monte Carlo Raises $16M to Build the World’s First Data Reliability Platform
- Talend Introduces Industry-First Measure of Data Health to Bring Clarity and Confidence to Every Business Decision
- IBM Cognos Analytics-Based Business Transformation Going Strong
- Scality RING8 on All-Flash Delivers File and Object Storage Performance 10x Faster Than Competitive Solutions
- ScyllaDB Unveils One-Step Migration from Amazon DynamoDB to Scylla NoSQL Database
- Tamr Data Mastering Platform Now Available on Microsoft Azure
- Kinetica Releases New Version of The Kinetica Streaming Data Warehouse Platform
- VMware and DataStax Partner to Bring Cloud-Native, Scale-Out, Hybrid Database-as-a-Service to Enterprises
- Spectra Logic Announces Industry’s First Tape Library to Store One Exabyte of Uncompressed Data Leveraging LTO-9 Technology
- AWS and the National Football League Announce New Next Gen Stats Powered by AWS for the 2020 Season
- More This Just In…