# Naive Bayes Classifier

I'm sure we've all had our fair share of encounters with probabilities, some people have it worse with conditional probability and I think Bayes theorem is when it becomes advanced. But to be honest, what is advanced in some basic "The probability of A given B is equals the probability of B given A multiplied by the probability of A all divided by the probability of B". Read this out loud for context:

P(A/B) = \[P(B/A) \* P(A)\]/P(B)

It is called the Bayes theorem. It is used in machine learning to predict target values by using the conditional probability of the features.

The naive part stems out of the assumption that all the features are actually without doubt, independent of each other as this strengthens the idea of the probabilities being independent and conditional rather than dependent. Even though this might not always be absolutely true, naivety ensures simplicity.

There are three types of naive bayes model to be used in ML. They are Bernoulli, Multinomial and Gaussian. Bernoulli is used when all the features are binary, 0 and 1, yes and no features only. Multinomial is used when the values of the features are discrete or completely categorical with different figures representing different categories. And Gaussian is used when the features contain continous values.

My mini project uses the properties of wine to classify the wine into different classes: class\_1, class\_2 and class\_3. I tested both the multinomial and gaussian naive bayes models to see which one worked best. I also used both the test train split method and the K-Fold cross validation technique to ensure I get more solid results.

![](https://media.licdn.com/dms/image/D4D12AQHPqWUbGdNoWA/article-inline_image-shrink_1500_2232/0/1701757162140?e=1707350400&v=beta&t=xbG8EBCUz3md7z6uFydjLgRSNZNvoPBXL8XNeQnDIQ0 align="left")

Part 1

![](https://media.licdn.com/dms/image/D4D12AQG6ihlXdX-Fhg/article-inline_image-shrink_1500_2232/0/1701757215518?e=1707350400&v=beta&t=hag56XJyixmZGb5Cw96uAc-3-IBR2ycOIVzDmlFj4kE align="left")

part 2

![](https://media.licdn.com/dms/image/D4D12AQG7yJRVqiRP0w/article-inline_image-shrink_1500_2232/0/1701757257146?e=1707350400&v=beta&t=WHuL6FBMegM9pZAqZWKzU-zoBu0FN8DxwP16LInEgao align="left")

part 3

![](https://media.licdn.com/dms/image/D4D12AQH8iQRd6O7Oiw/article-inline_image-shrink_1500_2232/0/1701757277613?e=1707350400&v=beta&t=zv4xjn91subZDvLpq-Pmx7GV3C45oYX15JAJs88GISY align="left")

part 4

![](https://media.licdn.com/dms/image/D4D12AQGHed2I0lAyFw/article-inline_image-shrink_1500_2232/0/1701757300012?e=1707350400&v=beta&t=jluU6YQmdneR1Po8P5vIeyUOKZ0fAowazOtvN3z04l8 align="left")

part 5

![](https://media.licdn.com/dms/image/D4D12AQEKjfWG_n6yoQ/article-inline_image-shrink_1500_2232/0/1701757331186?e=1707350400&v=beta&t=lx0gb7xQmKru8eophfgcqE_5-xBAY5pq7DjYPsUSp94 align="left")

Part 6 (The End)

As visible in the images of the workbook, Gaussian is superbly accurate. I would tell you that this is essentially due to the features like alcohol etc being continous values (check fig 2 above).

So, to use naive bayes......

```plaintext
from sklearn.naive_bayes import MultinomialNB
from sklearn.naive_bayes import GaussianNB
from sklearn.naive_bayes import BernoulliNB

MNB = MultinomialNB()
GNB = GaussianNB()
BNB = BernoulliNB()
```

So, yes bye.

ciao
