Tensors and reverse-mode autograd
Array operations build a dynamic graph. Backward traversal accumulates all contributions to a shared input and reduces broadcast gradients to their original shapes.
GitHub
Deep learning systemsA deep learning engine built from first principles, with its own gradients, trainable layers, GPT model, and native execution experiments.

Fluxion is a readable deep learning engine built from first principles. It owns its tensors, dynamic computation graph, backward rules, trainable layers and optimizers, then composes those pieces into a small GPT-style model. PyTorch is an independent numerical reference rather than the training engine underneath Fluxion.
Array operations build a dynamic graph. Backward traversal accumulates all contributions to a shared input and reduces broadcast gradients to their original shapes.
Linear, activation, normalization and embedding layers compose with losses, recursive parameter discovery, SGD and Adam.
Multi-head causal attention, learned positions and pre-norm residual blocks produce GPT logits for next-token training and generation.
NumPy is the default path. An opt-in C++/BLAS Linear operator and experimental standalone CUDA Linear kernels expose native boundaries.
Understanding a model's training behavior requires seeing how gradients accumulate, how shapes broadcast and where complete training steps spend time. Moving an operator into native code is only useful if the actual workload improves.
Tensor operations record parent inputs and local backward rules while producing a model output and loss.
The graph and derivatives remain inspectable. For y = x*x + x at x = 3, the shared x receives 3 + 3 + 1, producing gradient 7.
Validation copies the same inputs and parameters into independent PyTorch models and checks both forward values and backward gradients.
Benchmarks include forward, loss, backward and update. Operator-level work does not automatically translate into a faster model.
The October 4 Linux CPU/native checkpoint passed 72 tests with 2 CUDA skips. All five PyTorch reference reports passed. The retained workload measurements include 2,000 timed samples. The seeded tiny network's loss fell from 0.256 to 7.81e-13 after 1,000 updates.


Native Linear was slower on the measured Linux host. CUDA was not exercised in that checkpoint; device-resident GPT training is future work. The tiny repeated-text GPT example is a learning exercise, not evidence of broad language understanding.
A concise engineering brief and direct links to the implementation and supporting evidence.
A two-page project brief covering the purpose, workflow, design decisions, recorded checks, and limits.