Introduction
Data Science is being used in more and more areas and different industries to add value. Although the use cases and questions can vary greatly depending on the industry, there is a common process, the so-called CRISP-DM process, which is an approach for answering questions in a data driven manner. CRISP-DM stands for “Cross-Industrie Standard Process for Data Mining”, see this Wikipedia article for more details. In this blog, we will illustrate the CRISP-DM process and apply it the problem of predicting car prices based on certain features of cars, e.g. horsepower, mile per gallon (mpg) etc.
CRISP-DM
The CRISP-DM process consists of several steps which are summarized in the picture below:

- Business Understanding: The process starts with the Business Understanding. In that step, we define which problem we want to solve or which questions we want to answer.
- Data Understanding: The second step deals with the data which should be used to support the solution of the business problem. This data must be gathered and its meaning must be understood. For example, if the data is stored in a table, it is import to know the exact meaning of each column. The data understanding can make it necessary to get back to the first step, i.e. the business understanding.
- Data Preparation: In this step the data gets prepared for the next step, i.e. the data modelling. Preparing the data includes for example a strategy on how to deal with missing values. In general, the data is usually not that clean as one would expect it to be. But since data of low quality can lead to wrong answers to the business questions, the data preparation step includes also data cleaning.
- Modelling: This step is one of the core steps where one tries to find a model which describes the data. In terms of our business problem of predicting the price of a car based on its features, modellig means finding a function which maps the features of a car to its price. It could be possible that the modelling step requires to turn back to the data preparation step (for example if we realize in the modelling step that the data has not been sufficiently cleaned yet).
- Evaluation: Once the model is build, we need to check how good its performance is. This is done in the evaluation step. For example, we could predict the prices of cars based on the model for which we know the real prices and compare how close the predicted values are to the actual ones.
- Deployment: In the last step, if we are satisfied with the performance of the model, we need to productionize it. For the car price prediction example, that could mean that we build an application which allows to enter certain car features in a form and returns the prediction of the price to the user.
CRISP-DM Example Walkthrough for Car Prices
This section illustrates the process steps of CRISP-DM in terms of the car price prediction example. We only provide a high-level overview here, the technical details can be found in a GitHub Repository.
Business Problem
As mentioned above, the CRISP-DM process starts with the understanding of the business problem. Imagine for example a used car dealer who needs estimates what the price of a used car could be. The car dealer could be interest in predicting the price of a car based on its attributes. More precise, we try to answer to the following 3 business questions:
- Is the price of a car related to the horsepower?
- Is the price of a care related to the length of the car?
- Can the price of a car be predicted based in its attribute with reasonable accuracy?
Data Understanding and Data Preparation
The data is taken from the UCI machine learning repository and can be found here. It is important to read the documentation of the data which can be found using the same link as above, since the data itself has no column names. Below is a sample of the data:

The above sample does not contain all columns. We see attributes like ‘city-mpg’ and ‘horsepower’. The far right column is the price we want to predict based on the other attributes. The data contains missing values and hence data cleaning is necessary. We replace a mining value in a column by the average of all other values in that column.
Modelling
To come up with a model for predicting the price based on the other features, we use the k-nearest neighbors (kNN)algorithm. The idea behind kNN is simple: Given the features of a car (e.g. horsepower and city-mpg) for which we don’t know the price, we search for those k cars in the dataset which have the most similar values of this features and for which we know the price. Based on this k prices, we can then derive a prediction e.g. by taking the average of the k prices. More information on kNN can be found here.
Evaluation
To evaluate the performance of the kNN model, we use a metric called r2-score. The r2-score is the proportion of the variance in the target variable (which is the price in our example) that is predictable from the feature variables (which are e.g. horsepower and city-mpg). The r2-score is a number between 0 and 1 (or 0% to 100%). The closer the r2-score is to 1, the better the model. More details concerning the r2-score can be found here.
To answer the first two business questions, we derive a model which aims to predict the price by using just a single feature (we are particularly interested in the features horsepower and length since they relate to the first business questions) and evaluate the corresponding r2-scores. Here is a list of the results:

We see that horsepower has a r2-score of 76% which indicates that this features plays a role in predicting the price of a car. The length, on the other hand, does not play such a big role in explaining the price but
it is still in the upper half of features which are ordered w.r.t. r2-score.
Conclusion
Let us now come back to the business questions. We start with the first one:
- Is the price of a car related to the horsepower?
Based on the above table ‘features with r2-score’, we see that the feature horsepower has the highest r2-score (76%) among all considered features. Hence, it is definitively related to the price. Further the person correlation between the price and the feature horsepower is 81% which also indicates a clear relationship. The following scatterplot also displays a relationship:

Although the plot displays some heteroscedasticity, we can recognize a relationship between price and horsepower.
Now let us consider the second question:
- Is the price of a care related to the length of the car?
If we look at table ‘features with r2-score’, we see that the length is not among the top 5 features. Nevertheless, the person correlation between the price and the length is 69%. This indicates that also the length gives some hints in which price range a car could be. Also, the points in the following scatterplot are not complet randomly distributed:

Overall, we conclude that there exists also a relationship between the price and the length, although it is not that strong than the relationship between the price and the horsepower.
Let us now consider the last question:
- Can the price of a car be predicted based in its attribute with reasonable accuracy?
To answer the last business question, we build a model which uses more than one feature. Further, we optimize the model w.r.t. the k-value (i.e. how much neighbors should be use to derive the price?). The analysis (cf. the GitHub Repository mentioned above) shows, that the best model for car price prediction we could found is a kNN-model with k=4 and which uses just the two features horsepower and highway-mpg. With a r2-score of 85% we feel confident to give a positive answer to the last business question, i.e. we believe we ca predict the price of a care based on its features with reasonable accuracy.
Author: Dr. Christoph Berns, Data Scientist, areto consulting gmh