Lightweight, modular, and fully vectorized Deep Learning library in Python and NumPy. Designed for production: Keras-style API, numerically verified gradients, portable persistence without pickle, external caches for advanced architectures.
v1.0.0 — stable release. Package renamed to
nnlib, modern packaging withpyproject.toml, CI with linting, MIT LICENSE. See CHANGELOG.md for full history.
- No hidden mathematical coupling. Softmax implements the full Jacobian; CCE/BCE accept
from_logits=Truefor the stable shortcut. Nothing "silently assumes" who is before it. - Layers are pure w.r.t. trainable state.
forward(x) -> (output, cache),backward(d_output, cache) -> (d_input, grads). The same layer processes N inputs in parallel (siamese, triplet loss) without corruption. - Generic interface between layers and optimizers. Layers declare
parameters() -> Dict[str, ndarray]with any name/quantity. The optimizer doesn't know about hardcodedweights/biases. - Fail-fast on shapes.
build()propagates dimensions through the entire network atcompile()time, not on the firstfit(). - Portable persistence.
save(dir)producestopology.json(no pickle, readable, inspectable) +weights.npz(standard NumPy). Survives refactors.
Dense (alias Layer), Dropout, BatchNormalization.
Sigmoid, ReLU, LeakyReLU, ELU, Tanh, Softmax (full Jacobian), Linear.
SGD (with momentum and Nesterov), AdaGrad, RMSprop, Adam. All with clip_norm and clip_value.
MSE, MAE, Huber, BinaryCrossEntropy, CategoricalCrossEntropy, SparseCategoricalCrossEntropy. Cross-entropies accept from_logits.
HeNormal, XavierNormal, XavierUniform, Zeros, Ones.
L1, L2, L1L2 applicable to kernels.
BinaryAccuracy, CategoricalAccuracy, SparseCategoricalAccuracy, MeanAbsoluteError, RootMeanSquaredError, R2Score.
EarlyStopping, ModelCheckpoint, ReduceLROnPlateau, History.
- 68 tests covering layers, optimizers, losses, metrics, callbacks, persistence, state isolation, and integration.
- Numerical gradient check validating backprop including the Softmax+CCE path with real Jacobian.
- Siamese network demo verifying state isolation with shared weights.
git clone https://github.com/elJulioDev/Neural_Network.git
cd Neural_Network
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -e .For development (includes ruff and matplotlib):
pip install -e ".[dev]"importnumpyasnpfromnnlibimport (
NeuralNetwork, Dense, Dropout, BatchNormalization,
Adam, BinaryCrossEntropy, EarlyStopping,
)
X=np.array([[0,0],[0,1],[1,0],[1,1]], dtype=float)
y=np.array([[0],[1],[1],[0]], dtype=float)
model=NeuralNetwork()
model.add(Dense(8, input_size=2, activation='relu'))
model.add(Dense(1, activation='linear')) # logitsmodel.compile(
optimizer=Adam(learning_rate=0.05),
loss=BinaryCrossEntropy(from_logits=True), # stable path
)
model.fit(X, y, epochs=500, batch_size=4,
callbacks=[EarlyStopping(monitor='loss', patience=50)],
verbose=1)
logits=model.predict(X)
probs=1.0/ (1.0+np.exp(-logits))from_logits=True(recommended): final layerLinear, the loss applies softmax/sigmoid internally and uses the stable shortcut(pred - y) / N.from_logits=False: the previous layer can be any activation. Softmax propagates its full real Jacobian — mathematically correct with any loss.
Recommended production pattern:
# Binary classificationmodel.add(Dense(1, activation='linear'))
model.compile(loss=BinaryCrossEntropy(from_logits=True), ...)
# At inference:logits=model.predict(X)
probs=1/ (1+np.exp(-logits))
# Multiclass classificationmodel.add(Dense(n_classes, activation='linear'))
model.compile(loss=CategoricalCrossEntropy(from_logits=True), ...)
# At inference:logits=model.predict(X)
ex=np.exp(logits-logits.max(axis=1, keepdims=True))
probs=ex/ex.sum(axis=1, keepdims=True)model.save('my_model/')
# Produces:# my_model/topology.json <- architecture + optimizer + loss (readable)# my_model/weights.npz <- parameters + BatchNorm stateloaded=NeuralNetwork.load('my_model/')topology.json is inspectable, not executable, and survives internal refactors.
The same layer instance can process two different inputs without corruption. See examples/siamese_network.py.
fromnnlib.layerimportDenselayer=Dense(4, 3, activation='relu')
out1, cache1=layer.forward(x1)
out2, cache2=layer.forward(x2) # does NOT overwrite cache1d_in1, grads1=layer.backward(dL1, cache1) # uses cache1 — correctd_in2, grads2=layer.backward(dL2, cache2) # uses cache2 — correctDimensional mismatches are detected at build/compile time, not during training:
model=NeuralNetwork()
model.add(Dense(4, input_size=3, activation='relu'))
model.add(Dense(2, activation='softmax'))
model.compile(optimizer='adam', loss='cce')
# X has 10 features instead of 3 -> immediate ValueErrormodel.fit(np.random.randn(5, 10), ...)fromnnlibimport (
NeuralNetwork, Dense, Dropout, BatchNormalization,
Adam, L2, CategoricalCrossEntropy,
EarlyStopping, ReduceLROnPlateau,
)
model=NeuralNetwork()
model.add(Dense(64, input_size=20, activation='relu', kernel_regularizer=L2(0.001)))
model.add(BatchNormalization(64))
model.add(Dropout(0.3))
model.add(Dense(32, activation='relu'))
model.add(Dense(10, activation='linear')) # logitsmodel.compile(
optimizer=Adam(learning_rate=0.001, clip_norm=1.0),
loss=CategoricalCrossEntropy(from_logits=True),
)
model.fit(X_train, y_train, epochs=100, batch_size=32,
validation_split=0.2,
callbacks=[
EarlyStopping(monitor='val_loss', patience=15, restore_best_weights=True),
ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=5),
])
model.save('production_model/')fromnnlibimportNeuralNetworkimportnumpyasnpai_model=NeuralNetwork.load('/path/to/production_model/')
defpredict_view(request):
features=np.array([[...]]) # shape (1, n_features)logits=ai_model.predict(features)
ex=np.exp(logits-logits.max(axis=1, keepdims=True))
probs=ex/ex.sum(axis=1, keepdims=True)
returnJsonResponse({'probabilities': probs[0].tolist()})Neural_Network/
├── nnlib/ # Main package
│ ├── __init__.py # Public API re-exports
│ ├── activations.py # Stateless: forward -> (out, cache)
│ ├── callbacks.py
│ ├── initializers.py # With get_config for JSON
│ ├── layer.py # Dense, Dropout, BatchNorm; parameters() dict
│ ├── losses.py # from_logits in CCE/BCE
│ ├── metrics.py
│ ├── neural_network.py # External cache management, build(), save/load
│ ├── optimizers.py # Generic interface (layer_id, name, param, grad)
│ ├── regularizers.py
│ └── utils.py
├── tests/ # 68 tests
│ ├── test_activations.py # Stateless, Softmax Jacobian, config roundtrip
│ ├── test_gradient_check.py # Numerical backprop validation
│ ├── test_layer.py # Including state isolation test
│ ├── test_losses.py # Including from_logits path
│ ├── test_model.py # Integration + JSON+NPZ persistence
│ └── test_optimizers.py # Generic interface
├── examples/
│ ├── multiclass_classification.py
│ └── siamese_network.py # Demonstrates state isolation
├── main.py # XOR demo
├── pyproject.toml # Modern packaging (PEP 517/518)
├── requirements.txt # Dev dependencies
├── CHANGELOG.md
├── LICENSE # MIT
└── README.md
python -m unittest discover tests -vpython main.py
python examples/multiclass_classification.py
python examples/siamese_network.py- No convolutional or recurrent layers (Conv1D/2D, LSTM, GRU, etc.) — contributions welcome.
- No GPU acceleration — pure NumPy, CPU only.
- Full dataset must fit in memory — no streaming or lazy loading for large datasets.
- No built-in model export to ONNX or other formats.
MIT License. See LICENSE for details.