当前位置：网站首页>Machine learning practice - logistic regression-19

Machine learning practice - logistic regression-19

2022-07-28 12:49:00 【gemoumou】

Machine learning practice - Logical regression - User churn prediction

Insert picture description here

import numpy as np

train_data = np.genfromtxt('Churn-Modelling.csv',delimiter=',',dtype=np.str)
test_data = np.genfromtxt('Churn-Modelling-Test-Data.csv',delimiter=',',dtype=np.str)

x_train = train_data[1:,:-1]
y_train = train_data[1:,-1].astype(int)
x_test = test_data[1:,:-1]
y_test = test_data[1:,-1].astype(int)

x_train = np.delete(x_train,[0,1,2],axis=1)
x_test = np.delete(x_test,[0,1,2],axis=1)

x_train[:5]

Insert picture description here

y_train[:5]

Insert picture description here

# x_train[x_train=='Female'] = 0
# x_train[x_train=='Male'] = 1

from sklearn.preprocessing import LabelEncoder
labelencoder1 = LabelEncoder()
x_train[:,1] = labelencoder1.fit_transform(x_train[:,1])
x_test[:,1] = labelencoder1.transform(x_test[:,1])
labelencoder2 = LabelEncoder()
x_train[:,2] = labelencoder2.fit_transform(x_train[:,2])
x_test[:,2] = labelencoder2.transform(x_test[:,2])

Insert picture description here

x_train = x_train.astype(np.float32)
x_test = x_test.astype(np.float32)
y_train = y_train.astype(np.float32)
y_test = y_test.astype(np.float32)

from sklearn.preprocessing import StandardScaler
sc = StandardScaler()
x_train = sc.fit_transform(x_train)
x_test = sc.transform(x_test)

Insert picture description here

from sklearn.linear_model import LinearRegression
from sklearn.metrics import classification

LR = LinearRegression()
LR.fit(x_train,y_train)

predictions = LR.predict(x_test)
print(classification_report(y_test, predictions))

Insert picture description here

Machine learning practice - Logical regression - Diabetes prediction model

Insert picture description here

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

Insert picture description here

#  Load data 
diabetes_data = pd.read_csv('diabetes.csv')
diabetes_data.head()

Insert picture description here

#  Data and information 
diabetes_data.info(verbose=True)

Insert picture description here

#  Data description 
diabetes_data.describe()

Insert picture description here

#  Data shape 
diabetes_data.shape

Insert picture description here

#  View label distribution 
print(diabetes_data.Outcome.value_counts())
#  Use the histogram to draw the statistics of the number of labels 
p=diabetes_data.Outcome.value_counts().plot(kind="bar")
plt.show()

Insert picture description here

#  Visualizing data distribution 
p=sns.pairplot(diabetes_data, hue = 'Outcome')
plt.show()

Insert picture description here
There are mainly two types of pictures drawn here , Histogram and scatter . Histogram is used for single feature comparison , When comparing different features, scatter charts are used , Show the relationship between the two features . We can find some outliers by observing the data distribution , such as Glucose glucose ,BloodPressure Blood pressure ,SkinThickness Skin thickness ,Insulin Insulin ,BMI These characteristics of body mass index should be impossible 0 It's worth it .

#  Put glucose , Blood pressure , Skin thickness , Insulin , In body mass index 0 Replace with nan
colume = ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']
diabetes_data[colume] = diabetes_data[colume].replace(0,np.nan)

# pip install missingno
import missingno as msno
p=msno.bar(diabetes_data)
plt.show()

Insert picture description here

#  Set threshold 
thresh_count = diabetes_data.shape[0]*0.8
#  If the number of missing data in a column exceeds 20% It will be deleted 
diabetes_data = diabetes_data.dropna(thresh=thresh_count, axis=1)

p=msno.bar(diabetes_data)
plt.show()

Insert picture description here

#  Import interpolation Library 
from sklearn.preprocessing import Imputer 
#  Missing values for numeric variables , We use the mean interpolation method to fill in the missing values 
imr = Imputer(missing_values='NaN', strategy='mean', axis=0) 
colume =  ['Glucose', 'BloodPressure', 'BMI']
#  Interpolate 
diabetes_data[colume] = imr.fit_transform(diabetes_data[colume])

p=msno.bar(diabetes_data)
plt.show()

Insert picture description here

plt.figure(figsize=(12,10))  
#  Draw a heat map , The value is the correlation coefficient between the two variables 
p=sns.heatmap(diabetes_data.corr(), annot=True) 
plt.show()

Insert picture description here

#  Segment data into features x And labels y
x = diabetes_data.drop("Outcome",axis = 1)
y = diabetes_data.Outcome

from sklearn.model_selection import train_test_split
#  Sharding data sets ,stratify=y Represents the ratio of data types in the training set and test set after segmentation to that before segmentation y The proportion is the same 
#  For example, before segmentation y in 0 and 1 The proportion of 1:2, After cutting y_train and y_test in 0 and 1 The proportion is also 1:2
x_train,x_test,y_train,y_test = train_test_split(x,y,test_size=0.3, stratify=y)

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report

LR = LogisticRegression()
LR.fit(x_train,y_train)

predictions = LR.predict(x_test)
print(classification_report(y_test, predictions))