We’ve talked so much about data and its sheer volume. What follows data is dirty data. Dirty data is any incorrect data.
Susan Walsh, Founder & MD, of The Classification Guru, explains:
The interpretation of “incorrect” may differ from business to business. For instance, one company may categorize DHL as a “courier” while another may categorize it as “logistics” or “warehousing”. Hence, it is important to understand dirty data for your organization and accordingly cleanse and check data. ~ DigitalFirst
This article will explain dirty data with examples, analyze its impact on businesses, and explore ways to cleanse it.
What is Dirty Data?
Any data that is inaccurate, incomplete, or inconsistent is dirty data. To write a dirty data definition- it is data that is erroneous or cannot be easily used. Reports reveal that companies believe at least 26% of their data is dirty data, and it does incur enormous losses. Infact, it costs, on average, 15-16% of the revenue. Bad data costs US companies around three trillion dollars annually (IBM).
Anybody dealing with dirty data will tell you how frustrating and annoying it is, especially when you encounter errors like typos daily. You do not understand the impact of these small “errors” until the numbers add up and you see dirty data’s huge impact on businesses and their annual revenues.
Where does Dirty Data come from?
Dirty data has various forms and shapes. To understand it explicitly, you must understand what creates inaccurate data. Below are a few common reasons for inconsistent or inaccurate data –
-
Human error
The most significant cause of dirty data is human error, which contributes over 60%. This should be of no surprise as humans make a mistake and deal with data, and combining both aspects leads to the concept of dirty data.
-
Department miscommunication
Another significant cause of dirty data is inter-departmental communication or, rather, the lack of it. Poor communication between departments leads to 35% of dirty data.
This happens when one department provides data to another in an inefficient manner. Different departments often work in data silos that have their way of storing, handling, and formatting data; the moment data is moved from one department to the other, the problem of dirty data tends to arise.
-
Poor data strategy
Another reason for dirty data is the poor data strategy implemented by the organization in dealing with their data. If data is being input manually at any point, combining data has faults, or rudimentary tools are being used to access, store, or manage data; all these cause dirty data.
-
Customer disinterest or doubt
A lot of data is generated when the customers provide their information. However, the method of obtaining information matters a lot.
If the information is asked through a mail v/s is gained by asking someone when they are busy buying groceries, you can expect a difference in the data quality.
Also, getting correct data has various other factors. If the customer or the person providing data feels that the information can be misused, this doubt can lead them to give partial or wrong information.
-
Wrong form formats
An extension of the above-mentioned problem is when the survey forms are designed poorly. What I mean by poor design of the form is that you are asking a lot of detailed questions from the customer, or a lot of irrelevant information is being asked. This can make the customer give partial or wrong information.
Take a look at Grahan Lewis analyzing a badly formatted questionnaire. Take a cue to what to NOT write in a survey.