The Wayback Machine - https://web.archive.org/web/20240930205259/https://www.geeksforgeeks.org/python-statistics-variance/
Open In App

Python statistics | variance()

Last Updated : 29 Jul, 2024
Summarize
Comments
Improve
Suggest changes
Like Article
Like
Save
Share
Report
News Follow

Statistics module provides very powerful tools, which can be used to compute anything related to Statistics. variance() is one such function. This function helps to calculate the variance from a sample of data (sample is a subset of populated data). 
variance() function should only be used when variance of a sample needs to be calculated. There’s another function known as pvariance(), which is used to calculate the variance of an entire population.
In pure statistics, variance is the squared deviation of a variable from its mean. Basically, it measures the spread of random data in a set from its mean or median value. A low value for variance indicates that the data are clustered together and are not spread apart widely, whereas a high value would indicate that the data in the given set are much more spread apart from the average value. 
Variance is an important tool in the sciences, where statistical analysis of data is common. It is the square of standard deviation of the given data-set and is also known as second central moment of a distribution. It is usually represented by [Tex]s^{2}, \sigma ^{2}, \operatorname {Var} (X)  [/Tex]in pure Statistics.
Variance is calculated by the following formula : 
 

It’s calculated by mean of square minus square of mean
[Tex]\operatorname {Var} (X)=\operatorname {E} \left[(X-\mu )^{2}\right]  [/Tex]
 


 

Syntax : variance( [data], xbar )
Parameters : 
[data] : An iterable with real valued numbers. 
xbar (Optional) : Takes actual mean of data-set as value.
Returntype : Returns the actual variance of the values passed as parameter.
Exceptions : 
StatisticsError is raised for data-set less than 2-values passed as parameter. 
Throws impossible values when the value provided as xbar doesn’t match actual mean of the data-set. 
 


Code #1 :
 

Python3

# Python code to demonstrate the working of # variance() function of Statistics Module # Importing Statistics module import statistics # Creating a sample of data sample = [2.74, 1.23, 2.63, 2.22, 3, 1.98] # Prints variance of the sample set # Function will automatically calculate # it's mean and set it as xbar print("Variance of sample set is % s" %(statistics.variance(sample)))

Output : 
 

Variance of sample set is 0.40924


  
Code #2 : Demonstrates variance() on a range of data-types 
 

Python3

# Python code to demonstrate variance() # function on varying range of data-types # importing statistics module from statistics import variance # importing fractions as parameter values from fractions import Fraction as fr # tuple of a set of positive integers # numbers are spread apart but not very much sample1 = (1, 2, 5, 4, 8, 9, 12) # tuple of a set of negative integers sample2 = (-2, -4, -3, -1, -5, -6) # tuple of a set of positive and negative numbers # data-points are spread apart considerably sample3 = (-9, -1, -0, 2, 1, 3, 4, 19) # tuple of a set of fractional numbers sample4 = (fr(1, 2), fr(2, 3), fr(3, 4), fr(5, 6), fr(7, 8)) # tuple of a set of floating point values sample5 = (1.23, 1.45, 2.1, 2.2, 1.9) # Print the variance of each samples print("Variance of Sample1 is % s " %(variance(sample1))) print("Variance of Sample2 is % s " %(variance(sample2))) print("Variance of Sample3 is % s " %(variance(sample3))) print("Variance of Sample4 is % s " %(variance(sample4))) print("Variance of Sample5 is % s " %(variance(sample5)))

Output : 
 

Variance of Sample 1 is 15.80952380952381
Variance of Sample 2 is 3.5
Variance of Sample 3 is 61.125
Variance of Sample 4 is 1/45
Variance of Sample 5 is 0.17613000000000006


  
Code #3 : Demonstrates the use of xbar parameter 
 

Python3

# Python code to demonstrate # the use of xbar parameter # Importing statistics module import statistics # creating a sample list sample = (1, 1.3, 1.2, 1.9, 2.5, 2.2) # calculating the mean of sample set m = statistics.mean(sample) # calculating the variance of sample set print("Variance of Sample set is % s" %(statistics.variance(sample, xbar = m)))

Output : 
 

Variance of Sample set is 0.3656666666666667


  
Code #4 : Demonstrates the Error when value of xbar is not same as the mean/average value 
 

Python3

# Python code to demonstrate the error caused # when garbage value of xbar is entered # Importing statistics module import statistics # creating a sample list sample = (1, 1.3, 1.2, 1.9, 2.5, 2.2) # calculating the mean of sample set m = statistics.mean(sample) # Actual value of mean after calculation # comes out to 1.6833333333333333 # But to demonstrate xbar error let's enter # -100 as the value for xbar parameter print(statistics.variance(sample, xbar = -100))

Output : 
 

0.3656666666663053


Note : It is different in precision from the output in Code #3 
  
Code #4 : Demonstrates StatisticsError 
 

Python3

# Python code to demonstrate StatisticsError # importing Statistics module import statistics # creating an empty data-srt sample = [] # will raise Statistics Error print(statistics.variance(sample))

Output : 
 

Traceback (most recent call last):
File "/home/64bf6d80f158b65d2b75c894d03a7779.py", line 10, in
print(statistics.variance(sample))
File "/usr/lib/python3.5/statistics.py", line 555, in variance
raise StatisticsError('variance requires at least two data points')
statistics.StatisticsError: variance requires at least two data points


  
Applications : 
Variance is a very important tool in Statistics and handling huge amounts of data. Like, when the omniscient mean is unknown (sample mean) then variance is used as biased estimator. Real world observations like the value of increase and decrease of all shares of a company throughout the day cannot be all sets of possible observations. As such, variance is calculated from a finite set of data, although it won’t match when calculated taking the whole population into consideration, but still it will give the user an estimate which is enough to chalk out other calculations.
 



Similar Reads

sympy.stats.variance() function in Python
In mathematics, the variance is the way to check the difference between the actual value and any random input, i.e variance can be calculated as a squared difference of these two values. With the help of sympy.stats.variance() method, we can calculate the value of variance by using this method. Syntax : sympy.stats.variance(value) Return : Return t
1 min read
Find combined mean and variance of two series
Given two different series arr1[n] and arr2[m] of size n and m. The task is to find the mean and variance of combined series.Examples : Input : arr1[] = {3, 5, 1, 7, 8, 5} arr2[] = {5, 9, 7, 1, 5, 4, 7, 3} Output : Mean1: 4.83333 mean2: 5.125 StandardDeviation1: 5.47222 StandardDeviation2: 5.60938 Combined Mean: 5 d1_square: 0.0277777 d2_square: 0.
12 min read
Compute the mean, standard deviation, and variance of a given NumPy array
In NumPy, we can compute the mean, standard deviation, and variance of a given array along the second axis by two approaches first is by using inbuilt functions and second is by the formulas of the mean, standard deviation, and variance. Method 1: Using numpy.mean(), numpy.std(), numpy.var() Python Code import numpy as np # Original array array = n
2 min read
How to normalize a tensor to 0 mean and 1 variance in Pytorch?
A tensor in PyTorch is like a NumPy array with the difference that the tensors can utilize the power of GPU whereas arrays can't. To normalize a tensor, we transform the tensor such that the mean and standard deviation become 0 and 1 respectively. As we know that the variance is the square of the standard deviation so variance also becomes 1. We co
3 min read
Single Estimator Versus Bagging: Bias-Variance Decomposition in Scikit Learn
You can use the Bias-Variance decomposition to assess how well one estimator performs in comparison to the Bagging ensemble approach. We may examine the average predicted loss, average bias, and average variance for both strategies by using the bias_variance_decomp function. The bias-variance trade-off, wherein increasing model complexity decreases
6 min read
MANOVA (Multivariate Analysis of Variance)
A strong statistical method for evaluating the simultaneous effects of one or more independent variables on several dependent variables is a multivariate analysis of variance or MANOVA. We will look at how to use Python, a popular and flexible computer language for data analysis, for MANOVA in this tutorial. We'll go over the theoretical foundation
15+ min read
Variance and standard-deviation of a matrix
Prerequisite - Mean, Variance and Standard Deviation, Variance and Standard Deviation of an arrayGiven a matrix of size n*n. We have to calculate variance and standard-deviation of given matrix. Examples : Input : 1 2 3 4 5 6 6 6 6 Output : variance: 3 deviation: 1 Input : 1 2 3 4 5 6 7 8 9 Output : variance: 6 deviation: 2 Explanation: First mean
8 min read
Program to calculate Variance of first N Natural Numbers
Understanding variance is crucial in statistics as it measures the spread of data points around the mean. In this article, we explore how to calculate the variance of the first N natural numbers, denoted by symbol [Tex]\sigma^2[/Tex]. Calculating Variance of First N Natural NumbersGiven an integer N, the task is to find the variance of the first N
4 min read
Expression for mean and variance in a running stream
Calculating the mean and variance in a running data stream is a fundamental task in statistics and data analysis. These metrics provide insights into the central tendency and variability of data, respectively. Unlike static datasets, a running stream of data continuously updates, requiring efficient algorithms to compute these statistics in real ti
6 min read
Program for Variance and Standard Deviation of an array
Given an array, we need to calculate the variance and standard deviation of the elements of the array. Examples : Input : arr[] = [1, 2, 3, 4, 5]Output : Variance = 2 Standard Deviation = 1 Input : arr[] = [7, 7, 8, 8, 3]Output : Variance = 3 Standard Deviation = 1We have discussed program to find mean of an array. Mean is average of element. Mean
9 min read
Mean, Variance and Standard Deviation
Mean, Variance and Standard Deviation are fundamental concepts in statistics and engineering mathematics, essential for analyzing and interpreting data. These measures provide insights into data's central tendency, dispersion, and spread, which are crucial for making informed decisions in various engineering fields. This article explores the defini
8 min read
Use Pandas to Calculate Statistics in Python
Performing various complex statistical operations in python can be easily reduced to single line commands using pandas. We will discuss some of the most useful and common statistical operations in this post. We will be using the Titanic survival dataset to demonstrate such operations. C/C++ Code # Import Pandas Library import pandas as pd # Load Ti
6 min read
mode() function in Python statistics module
The mode of a set of data values is the value that appears most often. It is the value at which the data is most likely to be sampled. A mode of a continuous probability distribution is often considered to be any value x at which its probability density function has a local maximum value, so any peak is a mode.Python is very robust when it comes to
5 min read
median_low() function in Python statistics module
Median is often referred to as the robust measure of the central location and is less affected by the presence of outliers in data. statistics module in Python allows three options to deal with median / middle elements in a data set, which are median(), median_low() and median_high(). The low median is always a member of the data set. When the numb
4 min read
median_high() function in Python statistics module
Median is often referred to as the robust measure of the central location and is less affected by the presence of outliers in data. statistics module in Python allows three options to deal with median / middle elements in a data set, which are median(), median_low() and median_high(). The high median is always a member of the data set. When the num
4 min read
median_grouped() function in Python statistics module
Coming to Statistical functions, median of a data-set is the measure of robust central tendency, which is less affected by the presence of outliers in data. As seen previously, medians of an ungrouped data-set using median(), median_high(), median_low() functions. Python gives the option to calculate the median of grouped and continuous data functi
4 min read
stdev() method in Python statistics module
Statistics module in Python provides a function known as stdev() , which can be used to calculate the standard deviation. stdev() function only calculates standard deviation from a sample of data, rather than an entire population. To calculate standard deviation of an entire population, another function known as pstdev() is used. Standard Deviation
5 min read
Python - Moyal Distribution in Statistics
scipy.stats.moyal() is a Moyal continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0 scale : [op
2 min read
Python - Maxwell Distribution in Statistics
scipy.stats.maxwell() is a Maxwell (Pareto of the second kind) continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location pa
2 min read
Python - Lomax Distribution in Statistics
scipy.stats.lomax() is a Lomax (Pareto of the second kind) continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parame
2 min read
Python - Log Normal Distribution in Statistics
scipy.stats.lognorm() is a log-Normal continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0 scal
2 min read
Python - Log Laplace Distribution in Statistics
scipy.stats.loglaplace() is a log-Laplace continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0
2 min read
Python - Logistic Distribution in Statistics
scipy.stats.logistic() is a logistic (or Sech-squared) continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter.
2 min read
Python - Log Gamma Distribution in Statistics
scipy.stats.loggamma() is a log gamma continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0 scal
2 min read
Python - Levy_stable Distribution in Statistics
scipy.stats.levy_stable() is a Levy-stable continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0
2 min read
Python - Left-skewed Levy Distribution in Statistics
scipy.stats.levy_l() is a left-skewed Levy continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0
2 min read
Python - Laplace Distribution in Statistics
scipy.stats.laplace() is a Laplace continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0 scale :
2 min read
Python - Levy Distribution in Statistics
scipy.stats.levy() is a levy continuous random variable. It is inherited from the of generic methods as an instance of the rv_continuous class. It completes the methods with details specific for this particular distribution. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0 scale : [opti
2 min read
Python - Kolmogorov-Smirnov Distribution in Statistics
scipy.stats.kstwobign() is Kolmogorov-Smirnov two-sided test for large N test that is defined with a standard format and some shape parameters to complete its specification. It is a statistical test that measures the maximum absolute distance of the theoretical CDF from the empirical CDF. Parameters : q : lower and upper tail probability x : quanti
2 min read
Python - ksone Distribution in Statistics
scipy.stats.ksone() is a General Kolmogorov-Smirnov one-sided test that is defined with a standard format and some shape parameters to complete its specification. It is a statistical test for the finite sample size n. Parameters : q : lower and upper tail probability x : quantiles loc : [optional]location parameter. Default = 0 scale : [optional]sc
2 min read