SparseCoding is a small research-oriented Python project for learning sparse representations of data. It originates from master's internship work completed at IRIT in the Samova team, and the maintained code path in this repository trains a dictionary on the scikit-learn handwritten-digits dataset, infers sparse activation codes, and saves artifacts that can be inspected as a simple end-to-end demo.
The repository also contains several historical notebooks exploring traditional sparse coding, convolutional sparse coding, LC-KSVD, and speech experiments. Those notebooks are kept for reference, but they rely on older research dependencies and are not part of the supported demo workflow. A new supported Jupyter notebook demo is provided for a lightweight MNIST walkthrough.
- learns a dictionary of atoms from input samples
- infers sparse coefficients for reconstructing those samples
- visualizes the learned atoms as image tiles
- saves dictionary, code, and cost artifacts for inspection
- Python 3.10+
- pip
cd <repository-root>
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r Code/requirements.txtcd Code
python setup.py --namecd <repository-root>
python Code/demo.py --demo --output-dir demo_outputThe command writes these files into demo_output/:
dictionary.npy: learned dictionary matrixcodes.npy: sparse coefficient matrixcosts.npy: optimization cost historydictionary.png: image grid of learned atoms
- Install Jupyter if needed:
pip install notebook - From
<repository-root>, runjupyter notebook - Open
Code/MNIST Demo.ipynb
The notebook downloads MNIST from OpenML on first use, trains a compact dictionary on a small subset, and visualizes learned atoms plus reconstructions.
from SparseCoding import load_demo_digits, sparse_coding
samples = load_demo_digits(num_samples=16)
dictionary, codes, costs = sparse_coding(samples, k=8, return_costs=True)The demo CLI supports these options:
--output-dir: where demo files are written--samples: number of digit samples to use--atoms: number of dictionary atoms to learn--max-iter: number of alternating-optimization iterations--random-state: reproducible random seed
The sparse_coding() function also accepts:
alpha: ISTA learning ratelambda_coef: sparsity penaltyista_max_iter: inner sparse-code iterationstolerance: convergence thresholdreturn_costs: whether to return the cost history
Run the test suite with:
cd <repository-root>
python -m unittest discover -s tests -vThis code is made available for research and educational use. If you use this repository or build on this work, please cite the following paper:
@inproceedings{rolland2019label,
title={Label-consistent sparse auto-encoders},
author={Rolland, Thomas and Basarab, Adrian and Pellegrini, Thomas},
booktitle={Workshop on Signal Processing with Adaptative Sparse Structured Representations (SPARS 2019)},
pages={1--2},
year={2019}
}