PROJECT / 001 · MACHINE LEARNING
C Machine Learning Library
Machine learning, closer to the metal.
Conceptual illustration · not a screenshot or measured result
Overview
MLC is a completed C11 learning project covering matrix operations, dataset preparation, linear and logistic regression, and configurable dense neural networks. It uses the C standard and math libraries rather than an external ML framework, with end-to-end XOR and Iris training workflows in the test suite.
Motivation
Building the underlying pieces is a way to see what an abstraction hides: how numerical operations fit together and how training mechanisms translate into code.
How it works
Rows represent samples and columns represent features or outputs. Each dense layer computes XW + b, then applies its activation. Training runs a forward pass, evaluates a loss gradient, and propagates it backward through the layers before updating weights and biases with full-batch gradient descent. CSV loading, paired shuffling, train/test splitting, and standard scaling support the data pipeline.
Architecture
Public APIs live under include/mlc, with implementations in src. Matrix and dataset modules support the regression models; activation, loss, and dense-layer modules compose into the network API. CMake builds a static library, and CTest registers separate suites for the eight main modules. Fallible operations return MLCStatus codes, and callers release owned objects through matching free functions.
Implementation decisions
Matrices store double-precision values on the heap. Dense layers use seeded Xavier-uniform weights and zero biases for repeatable initialization. The loss supplies the averaging factor, so dense-layer backpropagation does not average gradients again. In the Iris workflow, the scaler is fitted only on training data, then reused on both partitions to avoid leaking test-set statistics.
Challenges
The engineering constraints include keeping matrix shapes consistent across forward and backward passes, managing temporary allocations, and rejecting invalid arguments or malformed data. Tests inspect gradient calculations, exact parameter updates, seeded behavior, and end-to-end training. The completed scope is deliberately bounded: dense networks and full-batch training, without mini-batches, convolutional layers, or model serialization.
What I learned
Implementing the training pipeline makes the relationship between the chain rule, matrix dimensions, and parameter updates concrete. It also connects numerical code with API design: ownership, error reporting, and reproducibility are part of making the mathematics usable. XOR and Iris provide small, understandable ways to exercise the complete pipeline rather than checking each operation only in isolation.