• Sign up
  • Log in
ABM – Machine Learning Platform for everyoneABM – Machine Learning Platform for everyoneABM – Machine Learning Platform for everyoneABM – Machine Learning Platform for everyone
  • How it works
  • Pricing
  • Developers
  • Blog
NextPrevious
correlation_causation_example1

Correlation does not imply causation

By Algolytics | Predictive models | Comments are Closed | 12 October, 2016 | 5

A popular phrase tossed around when we talk about statistical data is “there is correlation between variables”. However, many people wrongly consider this to be the equivalent of “there is causation between variables”. It’s important to explain the distinction:

Correlation means that once we know how one variable changes we can make reasonable deductions about how other variables change

There are several variants of correlation:

1. Positive correlation

Positive correlation means that with an increase/decrease of one variable, the other variable will rise/fall. A good analogy would be the number of stories in an apartment building and the number of apartments. Logically, a higher building will probably contain more apartments.

Full correlation is equal to 1 and means that we can use data to explicitly model changes in values of one variable when we know how others change. In practice this rarely happens, or it describes situations which are obvious and not useful. For example, an increase in Oil volume has perfect correlation with Oil mass.

correlation2

correlation1

2. Zero correlation

Correlation that equals 0 means there is no way to deduce behavior of one variable based on another variable’s behavior.

3. Negative correlation

Negative correlation means that with the rise/fall of one variable, we can expect a decrease/increase of the other variable.

Negative correlation that is above -1 means only a partial possibility of deduction. For example, if a person’s weight and running speed are negatively correlated, a heavier person can’t usually run as fast as a lighter person – but it’s not always the case.

correlation4

Full negative correlation equals -1 and means that we can perfectly deduce the fall/rise of one variable knowing the rise/fall of the other. In practice it rarely happens or is obvious and, therefore, not useful.

What does correlation mean?

Many people consider correlation as sure proof of causality between variables. That’s not true – correlation can be explained by 5 possibilities:

  1. Variable A influences Variable B – for example the combined salary in a household and the number of cars
  2. Variable B influences Variable A – as in the above example
  3. Variable A influences Variable B and Variable B influences Variable A – education level influences the wealth of a person and their wealth influences the education levels they can achieve.
  4. There exists an unknown variable C, which is correlated and influences both A and B – both the number of cars and the market price of the house in a household are influenced by the salaries.
  5. Correlation is a random accident.

It may seem that the last point is somehow lazy, but because of the sheer number of data available and the rules of probability, we can find large numbers of correlated variables that in practice are not connected at all.

Examples?

correlation_causation_example1

 

correlation_causation_example2The webpage http://www.tylervigen.com/spurious-correlations offers a great selection of these types of combinations.

Enjoy. 🙂


Want to read more news like this? Sign up for our Newsletter!

NAME

EMAIL

I agree to the processing of my personal data for the purpose of sending marketing information.

The administrator of the data given in the above form is Algolytics Technologies Sp. z o. o., ul. Przeskok 2, 00-032 Warszawa, NIP: 701-080-13-66, Regon: 369456263, District Court for the Capital City of Warsaw in Warsaw, XII Commercial Division of the National registered under KRS number 0000074723, Amount of the share capital: 321 300,00 PLN. Data is provided voluntarily and processed in order to respond to enquiries made using the form and to send marketing information. We would like to inform you about your right to be forgetten, your right to access the data and your right to correct it. Please note that your consent may be revoked at any time by sending an e-mail to gdpr@algolytics.pl from the address to which consent relates.

I accept Terms of service and Privacy policy

Read Terms of service

Read Privacy Policy

Share
Share3
Tweet
3 Shares
Predictive models

Algolytics

More posts by Algolytics

Related Post

  • Understanding machine learning #3: Confusion matrix – not all errors are equal

    By Algolytics | Comments are Closed

    One of the most typical tasks in machine learning is classification tasks. It may seem that evaluating the effectiveness of such a model is easy. Let’s assume that we have a model which, based onRead more

  • Understanding machine learning #2: Do we need machine learning at all?

    By Algolytics | Comments are Closed

    In the previous post of our Understanding machine learning series, we presented how machines learn through multiple experiences. We also explained how, in some cases, human beings are much better at interpreting data than machines.Read more

  • Understanding Machine Learning #1 – How machines learn?

    By Algolytics | Comments are Closed

    “If (there) was one thing all people took for granted, (it) was conviction that if you feed honest figures into a computer, honest figures (will) come out. Never doubted it myself till I met aRead more

  • How to assess quality and correctness of classification models? Part 4 – ROC Curve

    By Algolytics | Comments are Closed

    In the previous parts of our tutorial we discussed: Basic notation used in assessing classification models Quantitative quality indicators Confusion Matrix In this fourth part of the tutorial we will discuss the ROC curve. WhatRead more

  • Tutorial: How to establish quality and correctness of classification models? Part 3 – Confusion Matrix

    By Algolytics | Comments are Closed

    In the previous parts of the tutorial (part 1, part 2) we introduced quantitative indicators of classification model quality. In the next two parts we will take a closer look at a couple of graphicalRead more

NextPrevious

100px white

Created with love by Algolytics

+48 691 303 305
abm_support (at) algolytics.com

Company

  • About us
  • Blog
  • Contact Us

Product

  • Documentation
  • Pricing
  • Terms of service
  • Privacy policy
  • API
Copyright © 2020 Algolytics Technologies | All Rights Reserved
  • How it works
  • Pricing
  • Developers
  • Blog
ABM – Machine Learning Platform for everyone
This site uses cookies: Find out more.