Why Big Data and Data Scientists Are Overrated
What does it take to get value out of data? Many organizations assume that you need a big collection of data and a highly skilled data scientist to spin all those 1s and 0s into dollar signs. In reality, companies need neither of those things to be successful with data.
One of the biggest mistakes that organizations can make with their data analytics projects is to assume they need a data scientist at the very beginning. According to Daniel Mintz, chief data evangelist with Looker, organizations are much better off starting lower on the data analytics food chain and working their way up as they gain proficiency.
“I’ve seen cases where people hired a data scientist way before they’re ready,” Mintz tells Datanami. “They don’t actually have any data, and even if they do, it’s dirty and dispersed across a whole bunch of places. The data scientist who doesn’t necessarily understand their business arrives and says ‘Where’s the nicely curated data set that you want me to use to solve problems?’ And they say, ‘Oh we didn’t’ know that was a prerequisite.'”
The fact is, data scientists spend about three-quarters of their time doing data janitorial work – collecting, transforming, and cleaning data – rather than building the complex predictive models that they were actually hired for. That equals frustration for data scientists who had high hopes of making an impact, and sour grapes for the people who hired them.
Organizations should start with the basics, and work up from there. Instead of being lured by the “shiny object” syndrome and thinking you need a big Hadoop data lake or neural networks to solve a problem, seek the simplest answer.
“People make a mistake if they jump right to the most sophisticated tool, because they’re wasting a lot of time,” Mintz says. “The reality is a lot of problems are quite tractable with a simple regression. And some problems don’t even need that. You can just look at the data and see what’s happening.”
Mintz’s personnel advice? Hire data generalists who can do the time-consuming data legwork that’s needs to be done before more highly skilled (and highly paid) data scientists come in to do their highly specialized thing.
“The really key skill is having somebody who can take what is fundamentally a business question and translate that into a data question,” says Mintz, who previously at MoveOn.org and other data-intensive operations. “That’s the key skill. When you’re not big enough to have specialists, the business people, who aren’t data people, will know what the right business questions are.”
Mintz recommends pairing a SQL-loving analyst with an ETL-loving engineer to start helping the business prepare themselves to answer questions with data. As they document their data stores, define organization-specific metrics, and create workflows that transform and combine data in reliable and useful ways, they will start to see how the super powers of the real data scientists could best be used.
“As they scale up,” Mintz says, “they realize ‘Now we have three or five analysts, now we’re ready to add a data scientist, because now we’ve got a handle on what our data means, we know where things live in our schema, we know what problems might be tractable using a more sophisticated algorithms.'”
There are real benefits to be had by analyzing data, but like everything else in life, you must walk before you run. Hiring the right people to make up your data analytics team, and hiring them in the right order, is important.
“Folks are looking for a magical unicorn who can do it all, when in reality it’s a team sport,” Mintz says. “You really need to be thinking about how does the team work together, and how as a team do you cover all the bases. That starts with somebody who’s a utility player who can play all the positions, then you start to specialize.”
Follow the Data Crumb
Just as you don’t start with a data scientist, you shouldn’t start with big data, either. In fact, it’s much better to start with the right piece of data, however small that is.
For Wolf Ruzicka, the chairman of Washington D.C.-based analysis firm EastBanc Technologies, it starts with a single crumb of data.
“Just the other day I ran into a company that has accumulated 50PB of data,” Ruzicka tells Datanami. “That’s great. But when you compare them against competitors…all the metrics — profit margin, revenue, growth, size — really are the same. They were very proud of that 50PB of data lake. But really when I look at it, it must have turned into a data swamp.”
When EastBanc engages a new client, there is a flurry of activity and brain storming meetings as EastBanc analysts do their best to understand the business problem at issue, and the potential data available to solve it.
The company starts small and works quickly. The customer may have more pressing questions they want answered, but starting with the low-hanging fruit on easily explored data is a good way to get started, and provide validation that the analytics are worthy. Setting a hard initial deadline of two to four weeks helps encourage fast iteration.
That first piece of useful data becomes a “data crumb” that typically leads to further success, Ruzicka says. “That’s what we call it,” he says. “One data crumb of relevant data, and we iterate from there.”
When you draw it out on whiteboard, it looks very different than a typical big data architectural drawing. “It’s more of a data tree that you’re starting to groom,” he says. “You may end up with big data. But you don’t start with big data. You essentially turn it upside down.”
This approach is anathema to the current wave of big data thinking, which says one should throw all of one’s data into Hadoop, and hope that magical algorithms can make sense of it down the line. This approach may work, but most likely through sheer luck, Ruzicka says.
Ruzicka’s advice: It’s better to start with a smaller data set that’s more reliable and useful, than to start with a bigger data of unknown value.
“Instead of being this pathological data hoarder, rather be someone who assembles the data and continuously goes through data spring cleaning at very regular intervals,” he says. “Just as bad as it may be not to have any data, it’s just as bad, confusing and expensive to have lots and lots of data and not make any use of it.
“So why not find that middle ground, where you iterate around data breadcrumbs that have correlations with each other, that bring value to each other, and then your purposefully build up that big data base that you may ultimately end up with,” he continues. “Just find something of value and iterate from there, and over time you will answer the unknown unknowns that you were not even aware of in the beginning.”
October 18, 2021
- Fujitsu Analyzes Japanese Election Data with Foundry from Palantir Technologies
- WANdisco Announces General Availability of LiveData Platform for Azure
- Akridata Joins National Exascale Day Celebrations
October 15, 2021
- Elastic And Optimyze Join Forces to Deliver Continuous Profiling Platform
- Coveo Acquires Qubit
- Aicadium and SambaNova Partner to Bring AI Hardware Solution to Singapore
October 14, 2021
- Kinetica Now Accessible as a Service on Microsoft Azure
- Deloitte Launches CognitiveSpark for Marketing AI Solution
- Alation Acquires Artificial Intelligence Vendor Lyngo Analytics
- WeRide Relies on Alluxio for its Hybrid Cloud Storage Gateway for ML and AI
- FUJI Launches Sustainable Data Storage Initiative
- Logi Analytics Announces Logi Spark 2021 Virtual Conference
October 13, 2021
- Deephaven Community Core with Real-Time Data Capabilities Now Available
- Geospark Analytics Awarded Four-Year Contract from Department of State
- SparkBeyond Unveils No-Code AI Analytics Platform
- Dataminr is Acquiring Krizo, a Real-time Crisis Response Platform
- New Relic Launches Open Source Ecosystem of Quickstarts and Partner Integrations
- LogDNA Introduces Control API Suite to Give Customers More Control
- Elastic Announces Expanded Integrations with Google Cloud
- CrowdStrike Launches Free Humio Community Edition
Most Read Features
- Google Cloud Gives Spanner a PostgreSQL Interface
- One on One with Google Cloud Product Director Irina Farooq
- What Is Data Science? A Turing Award Winner Shares His View
- Big Data File Formats Demystified
- We’re In the Moneyball 3.0 Era. Here’s What It Means for Live Sports
- SambaNova Brings Custom Silicon To Bear on High-End AI Workloads
- Who’s Winning In the $17B AIOps and Observability Market
- What’s the Difference Between AI, ML, Deep Learning, and Active Learning?
- Five Real-World Applications for Sports Analytics
- How the Coronavirus Response Is Aided by Analytics
- More Features…
Most Read News In Brief
- Data and AI Salaries Continue Upward March, O’Reilly Says
- LinkedIn Open Sources Tech Behind 10,000-Node Hadoop Cluster
- Bigeye Observes $45 Million in Funding
- Data Prep Still Dominates Data Scientists’ Time, Survey Finds
- Gartner Shuffles the Technology Deck with Latest ‘Hype Cycle’ Report
- Why Is SAS Going Public?
- Feature Stores Emerging as Must-Have Tech for Machine Learning
- Sisu Nabs $62M to Grow Data Analytics Biz
- Logistics Operators Look to Data, Technology for Advantage
- An Interactive Analytics Whiteboard for COVID Times
- More News In Brief…
Most Read This Just In
- TIBCO NOW 2021 Showcases Limitless Power of Data
- Databricks Acquires Low-code/No-code Company to Expand its Lakehouse Platform
- Toloka Launches Data Research Grants, Announces First Eight Recipients
- BriefCam Introduces Video Analytics Enabled on Deep Learning Cameras from Axis Communications
- NetApp to Acquire CloudCheckr and Expand its Spot by NetApp CloudOps Platform
- Transaction Processing Performance Council (TPC) Launches an Artificial Intelligence Benchmark (TPCx-AI)
- Indico Data Announces General Availability of Indico Unstructured Data Platform
- Narmi Launches Narmi Analytics: Empowering Financial Institutions to Reclaim Control Over Data
- The Linux Foundation Announces Agenda and Speaker Lineup for the 2021 Linux Foundation Member Summit
- MicroAI to Bring AI Training to Renesas MCUs
- More This Just In…
Sponsored Partner Content
October 19London United Kingdom
October 27 - October 28
November 29 - December 3
December 6 - December 10San Diego CA United States
February 7, 2022 - February 9, 2022Houston TX United States
June 26, 2022 - June 30, 2022Hollywood FL United States