Predicting Horse Survival with Colic using Logistic Regression

A dataset for predicting survival in horses with colic contains 368 samples and 28 features. Colic refers to abdominal pain, which may stem from gastrointestinal or other issues. The dataset includes hospital measurements, some of which are subjective or difficult to quantify, such as pain level. Additionally, approximately 30% of the values are missing.

Handling Missing Data in the Dataset

Missing data presents significant challenges. For instance, if a sensor fails during data collection, rendering one feature invalid, disccarding the entire dataset is often impractical due to the cost of data acquisition. Several strategies exist to address missing values:

  1. Replace missing values with the mean of the available feature data.
  2. Use a specific placeholder value, such as -1.
  3. Omit samples containing missing values.
  4. Impute missing values using the mean from similar samples.
  5. Employ a machine learning algorithm to predict the missing values.

To prepare the data for logistic regression, two preprocessing steps are necessary:

  1. Replace all missing values with a real number. NumPy arrays cannot contain missing values. Using 0 is a suitable choice for logistic regression because the sigmoid function outputs 0.5 for an input of 0, indicating no bias in prediction. This approach preserves the existing data structure without altering the optimization algorithm's update rule: coefficients = coefficients + learning_rate * error * feature_vector. If a feature value is 0, its corresponding coefficient remains unchanged.
  2. Discard any test samples with a missing class label. Unlike features, it is difficult to reliably impute a missing target value.

The preprocessed data is saved in to two files: horseColicTraining.txt and horseColicTest.txt. The original dataset is available at the UCI Machine Learning Repository.

Implementing an Enhanced Stochastic Gradient Ascent Algorithm

Define the following functions in a Python module:

import numpy as np

def logistic(x):
    return 1.0 / (1 + np.exp(-x))

def stochastic_gradient_ascent(features, labels, iterations=150):
    m, n = features.shape
    coeff = np.ones(n)
    for epoch in range(iterations):
        index_list = list(range(m))
        for i in range(m):
            lr = 4.0 / (1.0 + epoch + i) + 0.01
            rand_idx = int(np.random.uniform(0, len(index_list)))
            selected_idx = index_list[rand_idx]
            h = logistic(np.sum(features[selected_idx] * coeff))
            error = labels[selected_idx] - h
            coeff += lr * error * features[selected_idx]
            del index_list[rand_idx]
    return coeff

Classsification with Logistic Regression

Classification involves multiplying the feature vector by the optimized coefficients, summing the result, and passing it through the logistic function. A result greater than 0.5 predicts class 1; otherwise, class 0.

Add the following functions to the module:

def predict(feature_vec, coefficients):
    prob = logistic(np.sum(feature_vec * coefficients))
    return 1.0 if prob > 0.5 else 0.0

def evaluate_model():
    train_file = open('horseColicTraining.txt')
    test_file = open('horseColicTest.txt')
    train_data = []
    train_labels = []
    for line in train_file.readlines():
        vals = line.strip().split('\t')
        feature_list = [float(vals[i]) for i in range(21)]
        train_data.append(feature_list)
        train_labels.append(float(vals[21]))
    model_coeff = stochastic_gradient_ascent(np.array(train_data), train_labels, 500)
    mistakes = 0
    total = 0.0
    for line in test_file.readlines():
        total += 1.0
        vals = line.strip().split('\t')
        feature_list = [float(vals[i]) for i in range(21)]
        if int(predict(np.array(feature_list), model_coeff)) != int(vals[21]):
            mistakes += 1
    error_rate = float(mistakes) / total
    print(f"Model error rate: {error_rate:.6f}")
    return error_rate

def run_multiple_tests(trials=10):
    total_error = 0.0
    for _ in range(trials):
        total_error += evaluate_model()
    avg_error = total_error / float(trials)
    print(f"Average error rate over {trials} trials: {avg_error:.6f}")

Executing the test yields results similar to the following:

>>> import horse_colic_model
>>> horse_colic_model.run_multiple_tests()
Model error rate: 0.343284
Model error rate: 0.283582
Model error rate: 0.432836
Model error rate: 0.358209
Model error rate: 0.328358
Model error rate: 0.328358
Model error rate: 0.343284
Model error rate: 0.373134
Model error rate: 0.432836
Model error rate: 0.462687
Average error rate over 10 trials: 0.368657

The average error rate is approximately 37%. Given the 30% rate of missing data, this performance is reasonable. The error rate can potentially be reduced to around 20% by tuning the number of iterations and the learning rate in the gradient ascent function.

Tags: Logistic Regression Machine Learning Data Preprocessing Missing Data Gradient Ascent

Posted on Thu, 08 Oct 2026 16:26:32 +0000 by magic-chef