Big data
Using big data to understand relationships between two data sets
Organizations increasingly rely on forecasting to plan ahead. One way to improve forecasts is to use big data - large, fast-moving data sets that can reveal patterns and relationships.
What is big data?
The Oxford dictionary defines it as, “extremely large data sets that may be analysed computationally to reveal patterns, trends, and associations, especially relating to human behaviour and interactions.”
In simpler terms, big data is information generated from the internet, social media, and other digital sources. It keeps growing as technology advances. Big data can take many forms, such as pictures, memos, videos, emojis, texts, gifs, and more.
Characteristics of big data
Velocity - the speed at which data is generated. If a business wants to collect this data, it needs systems that can capture and process it quickly.
If you use TikTok or Facebook, you may notice that even if you stay on the app all day, you keep seeing new content. That happens because new uploads arrive constantly.
According to the Social Shepherd website, over 272 videos are uploaded on TikTok per second, which leads to 16 000 videos per second and 985731 videos per hour. Now, imagine the numbers per day, per week, per month, and per year. This shows how quickly data can be generated.
Volume - the amount of data collected. As the example above suggests, data can arrive in extremely large quantities. Ordinary systems often can’t handle this because they aren’t designed to receive and store huge amounts of data at the same time.
Veracity - how truthful and reliable the data is. Many sources include fake news or misleading trends, and it’s not always easy to identify what’s true - especially as AI becomes more advanced. Organizations need to be careful when collecting and analysing data so they don’t base decisions on unreliable information.
Variety - the different forms data can take. Because big data includes many types (text, images, video, and more), a company may need tools and equipment that can organize the data into meaningful categories so it can be analysed.
The scatter graph
A scatter graph helps you explore the relationship between two variables by plotting paired data points on a graph. You then draw a line of best fit to see whether a linear relationship seems to exist.
A common point of confusion is the phrase “all variables.” A scatter graph doesn’t plot every variable in a data set at once. Instead, it uses all the data pairs for the two variables you’ve chosen.
Advantages of scatter graph
- Easy to use and to understand.
- It uses all the variables in the data set unlike just picking two like the high - low method does.
Disadvantages
- Just like high - low method it relies on historical data to predict the future.
- It also assumes that the activity level is the only factor that affect costs.
- As you have seen the line of best fit is a guess work, if I could have asked you to draw, you could have drawn a different one so it’s less reliable since it involves much guess work.
Regression analysis
The earlier methods have some clear weaknesses, especially when the line of best fit is drawn by eye. Regression analysis reduces this problem by calculating the line of best fit using a formula rather than guesswork.
The formula for finding the line of best fit is calculated as:
When n = number of pairs of data
And
The x is the arithmetic mean (or average) of x and is calculated as
The y is the arithmetic mean of y and is calculated as
The above formulas will give us the value of a, b, x and y.
The formula of analysing the semi-variable costs will be
Where:
- Y = total costs
- a = fixed costs
- b = variable cost per unit
- x = activity level