Selecting a Data Lake ETL Platform? Here Are 6 Questions to Ask
Not all data lakes are created equal. If your organization wants to adopt a data lake solution to simplify and more easily operate your IT infrastructure and store enormous quantities of data without requiring extended data transformation, then go for it.
But before you do, understand that simply dumping all your data into object storage such as AWS S3 doesn’t exactly mean you will have a working data lake.
The ability to use that data in analytics or machine learning requires converting that raw information into organized datasets you can use for SQL queries, and this can only be done via extract-transform-load (ETL) flows.
Data lake ETL platforms are available in a full range of options – from open-source to managed solutions to custom-built. Whichever tool you select, it’s important to differentiate data lake ETL challenges from traditional database ELT demands – and seek the platform that overcomes these obstacles.
Ask yourself which ETL solution:
1. Effectively Conducts Stateful Transformations
Traditional ETL frameworks allow for stateful operations like joins and aggregations to enable analysts to work with data from multiple sources; this is difficult to implement with a decoupled architecture.
Stateful transformations can occur by relying on extract-load-transform (ELT) – i.e., sending data to an “intermediary” database and using the database’s SQL, processing power and already amassed historical data. After transformation, the information is loaded into the data warehouse tables.
Data lakes, aiming to reduce cost and complexity by avoiding decoupled architecture, cannot depend on databases for every activity. You’ll need to look for an ETL tool that can conduct stateful transformations in-memory and needs no additional database to sustain joins and aggregations.
2 Extracts Schema from Raw Data
Organizations customarily use data lakes as a storehouse for raw data in a structured or semi-structured arrangement vs databases, which are predicated on structured tables. This poses numerous challenges.
One, in order to query data, can the data lake ETL tool draw out a schema (without which querying is not possible) from the raw data – and bring it up to date as changes in data and data structure come about? And two — this is an ongoing struggle — can the ETL tool effectively make queries with nested data?
3 Improves Query Performance Via Optimized Object Storage
Have you tried to read raw data straight from a data lake? Unlike using a database’s optimized file system that quickly sends back query results, doing the same operation with a data lake can be quite frustrating performance-wise.
To get optimal results, your ETL framework should continually store data in columnar formats and merge small files to the 200mb-1gb range. Unlike traditional ELT tools that only need to write the data once to its target database, data lake ETL should support the ability to write multiple copies of the same data based on the queries you will want to run and the various optimizations required for your query engines to be performant.
4 Easily Integrates with the Metadata Catalog
You’ve chosen the data lake approach for its flexibility — store large quantities if data now but analyze it later — and the ability to handle a wide range of analytics use cases. Such an open architecture should keep metadata separate from the engine that queries it, so you can easily change these query engines or use several simultaneously for the same data.
This means the data lake tool you select should reinforce this open architecture, i.e., be seamlessly merged with the metadata catalog. This allows the metadata to be easily “queryable” by various services because it is both stored in the catalog and still dovetails with every adjustment in schema, partition, and location of objects.
5 Replays Historical Data
Say you wanted to test a hypothesis by looking at stored data on a historical basis. This is difficult to accomplish with the traditional database option, where data is stored in a mutable condition, and in which running such a query could be prohibitive in terms of cost, stress, and tension between operations and analysis.
It’s easy to do with a data lake. In data lakes, stored raw data remains continuously available – it only transformed after extraction. Therefore, having a data lake allows you got “travel back in time,” seeing the exact state of the data as it was collected.
“Traditional” databases don’t allow for that, as the data is only stored in its transformed state.
6 Updates Tables Periodically
Data lakes, unlike databases that allow you to update and make deletions to tables, contain partitioned files that enable an append or add-only feature. If you want to store transactional data, implement change data capture in the data lake, or delete particular data for GDPR compliance, you’ll have difficulty doing so.
Make sure that the data lake ETL tools you choose have the ability to sidestep this obstacle. Your solution should be able to allow upserts, a system that lets you insert new records or update existing ones, in the storage layer and in the output tables.
About the author: Ori Rafael is the CEO and co-founder of Upsolver, a provider of a self-service data lake ETL platform that bridges the gap between data lakes and data consumers. Ori has worked in IT for nearly two decades and has an MBA from Tel Aviv University.
September 28, 2021
- Mitsui Chemicals Teams up With NEC and dotData to Trial AI-Based Price Change Forecasting
- Vertica Announces General Availability of Vertica Accelerator on AWS
- Pure Storage Unveils Portworx Data Services, the Industry’s First Database-as-a-Service Platform for Kubernetes
- Hazelcast Announces General Availability of Hazelcast Platform
- Oracle Introduces Next-Generation Exadata X9M Platforms
- Qlik Introduces Qlik Application Automation
- Fighting Fire with Data Science: UCSD Announces Joint Appointment with Los Alamos
September 27, 2021
- TIBCO Delivers a Comprehensive, Connected Platform for the Adaptable Digital Business
- The World Economic Forum Welcomes Western Digital to Global Lighthouse Network
- LevaData Introduces New Suite of Supply Management Software
- KNIME Data Talks: Bringing Business and Data Science Together; Set for September 29
- BriefCam Introduces Video Analytics Enabled on Deep Learning Cameras from Axis Communications
September 24, 2021
- AWS Announces General Availability of Amazon QuickSight Q
- IDC’s 3rd Platform Industry Spending Guides Provide In-Depth Sub-Industry Forecasts for Technology Investments Across Nine Industries
- Scality Awarded US Patent for Hyperscale Data Protection
September 23, 2021
- AtScale Expands Semantic Layer Solution for Microsoft Excel
- CNCF End User Technology Radar Provides Insights into DevSecOps
- At Annual OCEANS 2021, Sofar Ocean Debuts First-of-Its-Kind Maritime Open Standard, Bristlemouth
- Elastic Announces the General Availability of Elastic App Search Web Crawler, New Features for Elastic Enterprise Search
- Securonix Achieves FedRAMP In-Process Authorization
Most Read Features
- One on One with Google Cloud Product Director Irina Farooq
- Big Data File Formats Demystified
- What Is Data Science? A Turing Award Winner Shares His View
- Tabular Seeks to Remake Cloud Data Lakes in Iceberg’s Image
- What’s the Difference Between AI, ML, Deep Learning, and Active Learning?
- SambaNova Brings Custom Silicon To Bear on High-End AI Workloads
- Who’s Winning In the $17B AIOps and Observability Market
- How the Coronavirus Response Is Aided by Analytics
- In Search of the Modern Data Stack
- Rethinking Education in an AI-First World
- More Features…
Most Read News In Brief
- LinkedIn Open Sources Tech Behind 10,000-Node Hadoop Cluster
- Data and AI Salaries Continue Upward March, O’Reilly Says
- Data Prep Still Dominates Data Scientists’ Time, Survey Finds
- Gartner Shuffles the Technology Deck with Latest ‘Hype Cycle’ Report
- Who’s Winning in Open Source Data Tech
- Bigeye Observes $45 Million in Funding
- Why Is SAS Going Public?
- Hands-Off: Manual Data Integration Tasks Plummeting, Gartner Says
- Unstructured Data Growth Wearing Holes in IT Budgets
- Apollo CEO Bullish on GraphQL’s Potential in the Enterprise
- More News In Brief…
Most Read This Just In
- TIBCO NOW 2021 Showcases Limitless Power of Data
- Toloka Launches Data Research Grants, Announces First Eight Recipients
- Anaconda Announces Support for Pyston, Hiring Lead Developers Kevin Modzelewski and Marius Wachtler
- Kinetica Fuses Streaming and Contextual Analysis At Scale
- Transaction Processing Performance Council (TPC) Launches an Artificial Intelligence Benchmark (TPCx-AI)
- Aporia Launches Self-Serve Machine Learning Platform Open to Public
- MariaDB Announces SIS Provider Campus Cloud Services Migration to MariaDB SkySQL
- Snowflake Launches Financial Services Data Cloud
- OneTrust Enhances First-Party Data Solution to Strengthen Holistic Consent and Preference Management Platform
- Hewlett Packard Enterprise Wins $2B HPE GreenLake Contract with NSA
- More This Just In…
Sponsored Partner Content
October 5 - October 7
October 12 - October 14
October 19London United Kingdom
October 27 - October 28
November 29 - December 3
December 6 - December 10San Diego CA United States