Skip to content

Repository files navigation

PPC

DOIGitHub LicenseAsk DeepWiki

A model that performs binary classification to determine whether the subject is holding a smartphone. 48x48 RGB image.

class_idlabelModel output index
0no_possession0
1possession1

The PyTorch model and exported ONNX model always return two probabilities in the following order: [no_possession_probability, possession_probability].

output_.mp4
VariantSizeF1CPU
inference
latency
ONNX
P115 KB0.89590.24 msDownload
N176 KB0.93860.39 msDownload
T280 KB0.95200.51 msDownload
S495 KB0.96640.66 msDownload
C876 KB0.98680.73 msDownload
M1.7 MB0.99240.86 msDownload
L6.4 MB0.99611.07 msDownload

Data sample

no
possession
no
possession
possessionpossessionpossessionpossession
no_action_008364no_action_008001point_somewhere_002145point_somewhere_002068point_003496point_003008

Setup

The Python version and dependencies are pinned in pyproject.toml.

git clone https://github.com/PINTO0309/PPC.git &&cd PPC
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync
source .venv/bin/activate

Run the commands below in the managed environment by prefixing them with uv run.

Dataset

Place images under data/ using the following structure. Image filenames are not used to determine labels; only the top-level label directory is used.

data/
├── no_possession/
│ └── .../*.png
└── possession/
└── .../*.png

Generate the Parquet dataset. By default, raw image bytes are embedded, and each class is split into train and validation sets using a seeded 90/10 split.

uv run python 02_make_parquet.py

To store image paths without embedding image bytes, run:

uv run python 02_make_parquet.py --no-embed-images --overwrite

To merge multiple PPC Parquet files, specify each input path relative to data/.

uv run python 03_merge_parquet.py dataset_a.parquet dataset_b.parquet --overwrite

An optional preprocessing command can generate crops and annotations from videos or labeled image directories. Use --detector-model to specify the detector ONNX model.

uv run python 01_data_prep_realdata.py \
--input-image-dir /path/to/labeled-images \
--detector-model /path/to/detector.onnx
class_distribution
Split counts:
train: 39376
val: 4375
Label counts:
no_possession: 25981
possession: 17770
Split/label counts:
train no_possession: 23383
train possession: 15993
val no_possession: 2598
val possession: 1777

Inference

uv run python demo_phone_gaze_classification.py \
-v 0 \
-pm ppc_l_48x48.onnx \
-dlr -dnm -dgm -dhm \
-ep cuda \
-gm gazelle_dinov3_vit_tiny_inout_1x3x640x640_1xNx4.onnx \
--enable-heatmap
uv run python demo_phone_gaze_classification.py \
-v 0 \
-pm ppc_l_48x48.onnx \
-dlr -dnm -dgm -dhm \
-ep tensorrt \
-gm gazelle_dinov3_vit_tiny_inout_1x3x640x640_1xNx4.onnx \
--enable-heatmap

Training Pipeline

  • Use the labeled image folders under data/no_action, data/point_somewhere, and data/point.

  • 02_make_parquet.py writes pre-defined train/val splits into data/dataset.parquet using an image-level 9:1 split per class.

  • The training loop relies on BCEWithLogitsLoss plus class-balanced pos_weight to stabilise optimisation under class imbalance; inference produces sigmoid probabilities. Use --train_resampling weighted to switch on the previous WeightedRandomSampler behaviour, or --train_resampling balanced to physically duplicate minority classes before shuffling.

  • Training history, validation metrics, optional test predictions, checkpoints, configuration JSON, and ONNX exports are produced automatically.

  • Per-epoch checkpoints named like ppc_epoch_0001.pt are retained (latest 10), as well as the best checkpoints named ppc_best_epoch0004_f1_0.9321.pt (also latest 10).

  • The backbone can be switched with --arch_variant. Supported combinations with --head_variant are:

    --arch_variantDefault (--head_variant auto)Explicitly selectable headsRemarks
    baselineavgavg, avgmax_mlpWhen using transformer/mlp_mixer, you need to adjust the height and width of the feature map so that they are divisible by --token_mixer_grid (if left as is, an exception will occur during ONNX conversion or inference).
    inverted_seavgmax_mlpavg, avgmax_mlpWhen using transformer/mlp_mixer, it is necessary to adjust --token_mixer_grid as above.
    convnexttransformeravg, avgmax_mlp, transformer, mlp_mixerFor token mixer heads, the feature map dimensions must be divisible by --token_mixer_grid (default 2x3).
  • The classification head is selected with --head_variant (avg, avgmax_mlp, transformer, mlp_mixer, or auto which derives a sensible default from the backbone).

  • Pass --rgb_to_yuv_to_y to convert RGB crops to YUV, keep only the Y (luma) channel inside the network, and train a single-channel stem without modifying the dataloader.

  • Alternatively, use --rgb_to_lab or --rgb_to_luv to convert inputs to CIE Lab/Luv (3-channel) before the stem; these options are mutually exclusive with each other and with --rgb_to_yuv_to_y.

  • Mixed precision can be enabled with --use_amp when CUDA is available.

  • Resume training with --resume path/to/ppc_epoch_XXXX.pt; all optimiser/scheduler/AMP states and history are restored.

  • Loss/accuracy/F1 metrics are logged to TensorBoard under output_dir, and tqdm progress bars expose per-epoch progress for train/val/test loops.

Baseline depthwise-separable CNN:

SIZE=48x48
uv run python -m ppc train \
--data_root data/dataset.parquet \
--output_dir runs/ppc_${SIZE} \
--epochs 100 \
--batch_size 256 \
--train_resampling balanced \
--image_size ${SIZE} \
--base_channels 32 \
--num_blocks 4 \
--arch_variant baseline \
--seed 42 \
--device auto \
--use_amp

Inverted residual + SE variant (recommended for higher capacity):

SIZE=48x48
VAR=s
uv run python -m ppc train \
--data_root data/dataset.parquet \
--output_dir runs/ppc_is_${VAR}_${SIZE} \
--epochs 100 \
--batch_size 256 \
--train_resampling balanced \
--image_size ${SIZE} \
--base_channels 32 \
--num_blocks 4 \
--arch_variant inverted_se \
--head_variant avgmax_mlp \
--seed 42 \
--device auto \
--use_amp

ConvNeXt-style backbone with transformer head over pooled tokens:

SIZE=48x48
uv run python -m ppc train \
--data_root data/dataset.parquet \
--output_dir runs/ppc_convnext_${SIZE} \
--epochs 100 \
--batch_size 256 \
--train_resampling balanced \
--image_size ${SIZE} \
--base_channels 32 \
--num_blocks 4 \
--arch_variant convnext \
--head_variant transformer \
--token_mixer_grid 2x2 \
--seed 42 \
--device auto \
--use_amp
  • Outputs include the latest 10 ppc_epoch_*.pt, the latest 10 ppc_best_epochXXXX_f1_YYYY.pt (highest validation F1, or training F1 when no validation split), history.json, summary.json, optional test_predictions.csv, and train.log.
  • After every epoch a confusion matrix and ROC curve are saved under runs/ppc/diagnostics/<split>/confusion_<split>_epochXXXX.png and roc_<split>_epochXXXX.png.
  • --image_size accepts either a single integer for square crops (e.g. --image_size 48) or HEIGHTxWIDTH to resize non-square frames (e.g. --image_size 64x48).
  • Add --resume <checkpoint> to continue from an earlier epoch. Remember that --epochs indicates the desired total epoch count (e.g. resuming --epochs 40 after training to epoch 30 will run 10 additional epochs).
  • Launch TensorBoard with:
    tensorboard --logdir runs/ppc

ONNX Export

uv run python -m ppc exportonnx \
--checkpoint runs/ppc_is_s_48x48/ppc_best_epoch0049_f1_0.9939.pt \
--output ppc_s_48x48.onnx \
--opset 17
  • The saved graph exposes images as input and prob_pointing as output (batch dimension is dynamic); probabilities can be consumed directly.
  • After exporting, the tool runs onnxsim for simplification and rewrites any remaining BatchNormalization nodes into affine Mul/Add primitives. If simplification fails, a warning is emitted and the unsimplified model is preserved.

Arch

puc_p_48x48

Ultra-lightweight classification model series

  1. VSDLM: Visual-only speech detection driven by lip movements - MIT License
  2. OCEC: Open closed eyes classification. Ultra-fast wink and blink estimation model - MIT License
  3. PGC: Ultrafast pointing gesture classification - MIT License
  4. SC: Ultrafast sitting classification - MIT License
  5. PUC: Phone Usage Classifier is a three-class image classification pipeline for understanding how people interact with smartphones - MIT License
  6. HSC: Happy smile classifier - MIT License
  7. WHC: Waving Hand Classification - MIT License
  8. UHD: Ultra-lightweight human detection - MIT License
  9. MWC: Mask wearing classifier - MIT License
  10. SGC: Classification of wearing vs. not wearing sunglasses. 48x48. - MIT License
  11. HHC: Head Hat Classification. HHC is a binary classifier for cropped head images. 48x48. - MIT License
  12. BPC: Background Plain classification. 48x48. - MIT License
  13. PPC: Binary classification to determine whether the subject is holding a smartphone. 48x48 RGB image. - MIT License

Citation

If you find this project useful, please consider citing:

@software{hyodo2025ppc,
author = {Katsuya Hyodo},
title = {PINTO0309/PPC},
month = {07},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21422276},
url = {https://github.com/PINTO0309/ppc},
abstract = {Binary classification to determine whether the subject is holding a smartphone. 48x48 RGB image.},
}

About

Binary classification to determine whether the subject is holding a smartphone. 48x48 RGB image.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages