All projects

PROJECT / 001 · MACHINE LEARNING

C Machine Learning Library

Machine learning, closer to the metal.

Completed
  • C
  • Machine Learning
  • Numerical Computing

Conceptual illustration · not a screenshot or measured result

Overview

MLC is a completed C11 learning project covering matrix operations, dataset preparation, linear and logistic regression, and configurable dense neural networks. It uses the C standard and math libraries rather than an external ML framework, with end-to-end XOR and Iris training workflows in the test suite.

Motivation

Building the underlying pieces is a way to see what an abstraction hides: how numerical operations fit together and how training mechanisms translate into code.

How it works

Rows represent samples and columns represent features or outputs. Each dense layer computes XW + b, then applies its activation. Training runs a forward pass, evaluates a loss gradient, and propagates it backward through the layers before updating weights and biases with full-batch gradient descent. CSV loading, paired shuffling, train/test splitting, and standard scaling support the data pipeline.

Architecture

Public APIs live under include/mlc, with implementations in src. Matrix and dataset modules support the regression models; activation, loss, and dense-layer modules compose into the network API. CMake builds a static library, and CTest registers separate suites for the eight main modules. Fallible operations return MLCStatus codes, and callers release owned objects through matching free functions.

Implementation decisions

Matrices store double-precision values on the heap. Dense layers use seeded Xavier-uniform weights and zero biases for repeatable initialization. The loss supplies the averaging factor, so dense-layer backpropagation does not average gradients again. In the Iris workflow, the scaler is fitted only on training data, then reused on both partitions to avoid leaking test-set statistics.

Challenges

The engineering constraints include keeping matrix shapes consistent across forward and backward passes, managing temporary allocations, and rejecting invalid arguments or malformed data. Tests inspect gradient calculations, exact parameter updates, seeded behavior, and end-to-end training. The completed scope is deliberately bounded: dense networks and full-batch training, without mini-batches, convolutional layers, or model serialization.

What I learned

Implementing the training pipeline makes the relationship between the chain rule, matrix dimensions, and parameter updates concrete. It also connects numerical code with API design: ownership, error reporting, and reproducibility are part of making the mathematics usable. XOR and Iris provide small, understandable ways to exercise the complete pipeline rather than checking each operation only in isolation.

Repository

GitHub repository