All projects/ Fluxion02 / 05
Fluxion logoDeep learning systems

Fluxion.

Build the machinery behind a neural network, then measure it.

A deep learning engine built from first principles, with its own gradients, trainable layers, GPT model, and native execution experiments.

FLUXION / MEASURED CPU RUN
Measured Fluxion GPT CPU training-step latency and throughput across sequence lengths
72 tests passedRECORDED CHECKPOINT
THE IDEA

What Fluxion does.

Fluxion is a readable deep learning engine built from first principles. It owns its tensors, dynamic computation graph, backward rules, trainable layers and optimizers, then composes those pieces into a small GPT-style model. PyTorch is an independent numerical reference rather than the training engine underneath Fluxion.

01

Tensors and reverse-mode autograd

Array operations build a dynamic graph. Backward traversal accumulates all contributions to a shared input and reduces broadcast gradients to their original shapes.

02

Trainable modules

Linear, activation, normalization and embedding layers compose with losses, recursive parameter discovery, SGD and Adam.

03

A small causal Transformer

Multi-head causal attention, learned positions and pre-norm residual blocks produce GPT logits for next-token training and generation.

04

Explicit backend experiments

NumPy is the default path. An opt-in C++/BLAS Linear operator and experimental standalone CUDA Linear kernels expose native boundaries.

BUILT WITH

The technology.

FROM INPUT TO OUTCOME

How it works.

Understanding a model's training behavior requires seeing how gradients accumulate, how shapes broadcast and where complete training steps spend time. Moving an operator into native code is only useful if the actual workload improves.

01 / 04

Forward

Tensor operations record parent inputs and local backward rules while producing a model output and loss.

ENGINEERING DECISIONS

Why it’s built this way.

01

Readable correctness first

The graph and derivatives remain inspectable. For y = x*x + x at x = 3, the shared x receives 3 + 3 + 1, producing gradient 7.

02

Reference both directions

Validation copies the same inputs and parameters into independent PyTorch models and checks both forward values and backward gradients.

03

Measure the whole step

Benchmarks include forward, loss, backward and update. Operator-level work does not automatically translate into a faster model.

RECORDED CHECKPOINTS

What the evidence shows.

The October 4 Linux CPU/native checkpoint passed 72 tests with 2 CUDA skips. All five PyTorch reference reports passed. The retained workload measurements include 2,000 timed samples. The seeded tiny network's loss fell from 0.256 to 7.81e-13 after 1,000 updates.

72CPU/native tests passed
5 / 5PyTorch reference reports
2,000Retained timing samples
Measured Fluxion GPT CPU training-step latency and throughput across sequence lengths
Complete training-step measurements from the October 4 Linux CPU run.
For y = x*x + x at x = 3, shared-input gradient contributions 3 + 3 + 1 sum to 7.
Shared-input gradient accumulation, with a derivative you can check.
SCOPE & LIMITS

Native Linear was slower on the measured Linux host. CUDA was not exercised in that checkpoint; device-resident GPT training is future work. The tiny repeated-text GPT example is a learning exercise, not evidence of broad language understanding.

Read the engineering brief PDF
OPEN THE WORK

Code, documents & proof.

A concise engineering brief and direct links to the implementation and supporting evidence.

ENGINEERING BRIEFFluxionMichael Baffour Awuah · 2026

Inside the engineering

A two-page project brief covering the purpose, workflow, design decisions, recorded checks, and limits.