The Wayback Machine - https://web.archive.org/web/20240927174247/https://www.geeksforgeeks.org/ml-boston-housing-kaggle-challenge-with-linear-regression/
Open In App

ML | Boston Housing Kaggle Challenge with Linear Regression

Last Updated : 02 Aug, 2022
Summarize
Comments
Improve
Suggest changes
Like Article
Like
Save
Share
Report
News Follow

Boston Housing Data: This dataset was taken from the StatLib library and is maintained by Carnegie Mellon University. This dataset concerns the housing prices in the housing city of Boston. The dataset provided has 506 instances with 13 features.
The Description of the dataset is taken from the below reference as shown in the table follows: 

Let’s make the Linear Regression Model, predicting housing prices by Inputting Libraries and datasets.  

Python3




# Importing Libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
  
# Importing Data
from sklearn.datasets import load_boston
boston = load_boston()


The shape of input Boston data and getting feature_names. 

Python3




boston.data.shape


Python3




boston.feature_names


 Converting data from nd-array to data frame and adding feature names to the data 

Python3




data = pd.DataFrame(boston.data)
data.columns = boston.feature_names
 
data.head(10)


  Adding the ‘Price’ column to the dataset 

Python3




# Adding 'Price' (target) column to the data
boston.target.shape


Python3




data['Price'] = boston.target
data.head()


 Description of Boston dataset 

Python3




data.describe()


Info of Boston Dataset 

Python3




data.info()


Getting input and output data and further splitting data to training and testing dataset. 

Python3




# Input Data
x = boston.data
  
# Output Data
y = boston.target
  
  
# splitting data to training and testing dataset.
 
#from sklearn.cross_validation import train_test_split
#the submodule cross_validation is renamed and deprecated to model_selection
from sklearn.model_selection import train_test_split
 
xtrain, xtest, ytrain, ytest = train_test_split(x, y, test_size =0.2,
                                                    random_state = 0)
  
print("xtrain shape : ", xtrain.shape)
print("xtest shape  : ", xtest.shape)
print("ytrain shape : ", ytrain.shape)
print("ytest shape  : ", ytest.shape)


Applying Linear Regression Model to the dataset and predicting the prices.  

Python3




# Fitting Multi Linear regression model to training model
from sklearn.linear_model import LinearRegression
regressor = LinearRegression()
regressor.fit(xtrain, ytrain)
  
# predicting the test set results
y_pred = regressor.predict(xtest)


Plotting Scatter graph to show the prediction results – ‘y_true’ value vs ‘y_pred’ value.

Python3




# Plotting Scatter graph to show the prediction
# results - 'ytrue' value vs 'y_pred' value
plt.scatter(ytest, y_pred, c = 'green')
plt.xlabel("Price: in $1000's")
plt.ylabel("Predicted value")
plt.title("True value vs predicted value : Linear Regression")
plt.show()


Results of Linear Regression i.e. Mean Squared Error and Mean Absolute Error. 

Python3




from sklearn.metrics import mean_squared_error, mean_absolute_error
mse = mean_squared_error(ytest, y_pred)
mae = mean_absolute_error(ytest,y_pred)
print("Mean Square Error : ", mse)
print("Mean Absolute Error : ", mae)


Mean Square Error :  33.448979997676496
Mean Absolute Error :  3.8429092204444966

As per the result, our model is only 66.55% accurate. So, the prepared model is not very good for predicting housing prices. One can improve the prediction results using many other possible machine learning algorithms and techniques. 

Here are a few further steps on how you can improve your model.

  1. Feature Selection
  2. Cross-Validation
  3. Hyperparameter Tuning


Previous Article
Next Article

Similar Reads

Multiple linear regression analysis of Boston Housing Dataset using R
In this article, we are going to perform multiple linear regression analyses on the Boston Housing dataset using the R programming language. What is Multiple Linear Regression?Multiple Linear Regression is a supervised learning model, which is an extension of simple linear regression, where instead of just one independent variable, we have multiple
13 min read
Boston Housing Prices Datasets - Tensorflow.keras Datasets
Are you considering investing in the Boston real estate market? If so, you'll want to make informed decisions based on accurate data. That's where the Boston Housing Prices Datasets come into play. This comprehensive collection of datasets provides you with essential information on housing prices in the Boston area, allowing you to analyze trends,
7 min read
How to load Boston Housing data in python?
The Boston Housing dataset, which is used in regression analysis, provides insights into the housing values in the suburbs of Boston. This dataset has been a staple for algorithm demonstration, from simple linear regression to more complex machine learning models in predictive analytics. In this article, we will see how we can load the What is the
5 min read
Multiple Linear Regression using R to predict housing prices
Predicting housing prices is a common task in the field of data science and statistics. Multiple Linear Regression is a valuable tool for this purpose as it allows you to model the relationship between multiple independent variables and a dependent variable, such as housing prices. In this article, we'll walk you through the process of performing M
12 min read
Regression Models for California Housing Price Prediction
In this article, we will build a machine-learning model that predicts the median housing price using the California housing price dataset from the StatLib repository. The dataset is based on the 1990 California census and has metrics. It is a supervised learning task (labeled training) because each instance has an expected output (median housing pr
15+ min read
ML | Kaggle Breast Cancer Wisconsin Diagnosis using Logistic Regression
Dataset:It is given by Kaggle from UCI Machine Learning Repository, in one of its challenge. It is a dataset of Breast Cancer patients with Malignant and Benign tumor. Logistic Regression is used to predict whether the given patient is having Malignant or Benign tumor based on the attributes in the given dataset. Code : Loading Libraries [GFGTABS]
5 min read
ML | Linear Regression vs Logistic Regression
Linear Regression is a machine learning algorithm based on supervised regression algorithm. Regression models a target prediction value based on independent variables. It is mostly used for finding out the relationship between variables and forecasting. Different regression models differ based on – the kind of relationship between the dependent and
3 min read
The Difference between Linear Regression and Nonlinear Regression Models
areRegression analysis is a fundamental tool in statistical modelling used to understand the relationship between a dependent variable and one or more independent variables. Two primary types of regression models are linear regression and nonlinear regression. This article delves into the key differences between these models, their applications, an
7 min read
Support Vector Regression (SVR) using Linear and Non-Linear Kernels in Scikit Learn
Support vector regression (SVR) is a type of support vector machine (SVM) that is used for regression tasks. It tries to find a function that best predicts the continuous output value for a given input value. SVR can use both linear and non-linear kernels. A linear kernel is a simple dot product between two input vectors, while a non-linear kernel
5 min read
Boston Dataset in Sklearn
In this article, we are going to see how to use Boston Datasets using Sklearn. The Boston Housing dataset, one of the most widely recognized datasets in the field of machine learning, is a collection of data derived from the Boston Standard Metropolitan Statistical Area (SMSA) in the 1970s. This dataset is commonly used in regression analysis to pr
4 min read
Analyzing Housing Market Trends in R
Understanding housing market trends is crucial for various stakeholders, including investors, policymakers, real estate professionals, and homebuyers. Analyzing these trends involves examining factors such as housing prices, sales volume, interest rates, and economic indicators. R, with its powerful data manipulation and visualization capabilities,
6 min read
How to improve the performance of segmented regression using quantile regression in R?
Segmented regression, also known as piecewise or broken-line regression is a powerful statistical technique used to identify changes in the relationship between a dependent variable and one or more independent variables. Quantile regression, on the other hand, estimates the conditional quantiles of a response variable distribution in the linear mod
8 min read
Importing Kaggle dataset into google colaboratory
While building a Deep Learning model, the first task is to import datasets online and this task proves to be very hectic sometimes. We can easily import Kaggle datasets in just a few steps: Code: Importing CIFAR 10 dataset !pip install kaggle Now go to your Kaggle account and create new API token from my account section, a kaggle.json file will be
1 min read
Is Kaggle Useful in Finding a Machine Learning Job?
If you are in any way connected to the tech industry, chances are that you have heard of Machine Learning! It's the current cutting edge technology that has a wide range of applications in almost all sectors. And if you have heard of Machine Learning, chances are that you may be interested in learning more about it and enhancing your knowledge. Kag
5 min read
How Should a Machine Learning Beginner Get Started on Kaggle?
Are you fascinated by Data Science? Do you think Machine Learning is fun? Do you want to learn more about these fields but aren’t sure where to start? Well, start with Kaggle! Kaggle is an online community devoted to Data Scientist and Machine Learning founded by Google in 2010. It is the largest data community in the world with members ranging fro
8 min read
Getting started with Kaggle : A quick guide for beginners
Kaggle is an online community of Data Scientists and Machine Learning Engineers which is owned by Google. A general feeling of beginners in the field of Machine Learning and Data Science towards the website is of hesitance. This feeling mainly arises because of the misconceptions that the outside people have about the website. Here are some of them
3 min read
ML | Kaggle Breast Cancer Wisconsin Diagnosis using KNN and Cross Validation
Dataset : It is given by Kaggle from UCI Machine Learning Repository, in one of its challenges. It is a dataset of Breast Cancer patients with Malignant and Benign tumor. K-nearest neighbour algorithm is used to predict whether is patient is having cancer (Malignant tumour) or not (Benign tumour). Implementation of KNN algorithm for classification.
3 min read
how to switch ON the GPU in Kaggle Kernel?
The Kaggle is an excellent platform for data scientists and machine learning enthusiasts to experiment with models, datasets and tools. One of the powerful features of the Kaggle Kernels is the ability to leverage a GPU (Graphics Processing Unit) for computational tasks. This article will guide us through the steps to enable and verify GPU usage in
4 min read
7 Best Kaggle Kernels Alternatives in 2024
The Kaggle Kernels (now known as Kaggle Notebooks) have been a popular choice for data scientists and machine learning enthusiasts to create, share and collaborate on the code. However, there are several other platforms available in 2024 that offer similar functionalities and may even provide unique features or advantages. This article explores 7 B
8 min read
What is Kaggle?
Kaggle is a powerful online platform where the data science and machine learning community comes together. Launched in 2010, Kaggle provides a place in which data scientists, analysts, and machine learning enthusiasts can work on real-world problems, share knowledge, and participate in competitions. It brings a huge resource: datasets, code noteboo
8 min read
How to Use TPU in Kaggle
Tensor processing units, or TPUs, are becoming essential tools for machine learning engineers and data scientists. The processing capability of these customized chips is unmatched, greatly speeding up the processes of model training and inference. Effective use of TPUs allows you to experiment with various hyperparameters and investigate more intri
3 min read
How to Install PyPDF2 in Kaggle
Kaggle is a popular platform for data science and machine learning competitions and projects, providing a cloud-based environment with a range of pre-installed packages. However, there might be instances where you need additional libraries that aren't included by default. PyPDF2 is one such library for Python, used for working with PDF files — whet
3 min read
How to Install Twisted in Kaggle
Kaggle is a popular platform for data science and machine learning competitions, offering a collaborative environment with a variety of pre-installed libraries and tools. However, sometimes you may need to use libraries that aren't included by default in Kaggle's kernels. One such library is Twisted, an event-driven networking engine written in Pyt
3 min read
How to Install XlsxWriter in Kaggle
Kaggle provides a powerful platform for data science and machine learning, including support for various Python libraries. XlsxWriter is a popular library for creating and writing Excel files (.xlsx format). Follow these steps to install and use XlsxWriter in your Kaggle notebooks. Step 1: Open Your Kaggle NotebookGo to Kaggle: Visit Kaggle and log
3 min read
How to Install Openpyxl in Kaggle
Kaggle is a powerful platform for data science and machine learning, providing an environment to develop and execute Python code efficiently. The openpyxl library is a versatile tool for working with Excel files (.xlsx format). This guide will walk you through the process of installing and using openpyxl in your Kaggle notebooks. Step 1: Open a Kag
3 min read
How to Install PyYAML in Kaggle
Kaggle is a popular platform for data science and machine learning, providing a range of tools and datasets for data analysis and model building. If you're working on a Kaggle notebook and need to use PyYAML, a Python library for parsing and writing YAML, follow this step-by-step guide to get it up and running in your Kaggle environment. Step 1: Op
2 min read
How to Use GPU in Kaggle?
Using a GPU in Kaggle is simple and useful for deep learning or other computationally intensive tasks. Here's how you can enable and use a GPU in Kaggle: Steps to Enable GPU in Kaggle:Create or Open a Kaggle Notebook:Go to Kaggle Notebooks and create a new notebook or open an existing one.Enable GPU in the Notebook:Click on the “Settings” tab on th
2 min read
How to Install Scapy in Kaggle
If we’re working in a Kaggle notebook and want to use the Scapy library for network packet manipulation and analysis, you might be wondering how to install it. Installing Scapy in Kaggle is straightforward, and this article will walk you through the steps in simple language. To install Scapy in a Kaggle notebook, we can use the following steps. Kag
2 min read
How to Install Pandas-Profiling in Kaggle
If you’re working with data in Kaggle and want to quickly generate a detailed report on your dataset, Pandas-Profiling is a great tool to use. It helps you understand your data better by providing a comprehensive summary, including statistics, correlations, and missing values. Here's a simple guide on how to install and use Pandas-Profiling in a Ka
2 min read
How to Install AllenNLP in Kaggle
AllenNLP is a popular open-source library for natural language processing (NLP) developed by the Allen Institute for AI. It provides flexible tools for building deep learning models for various NLP tasks such as text classification, machine translation, and question-answering. One of the most notable features of AllenNLP is its ability to prototype
4 min read