In this paper, we present a practical deep learning (DL) approach for energy-efficient traffic classification (TC) on resource-limited microcontrollers, which are widely used in IoT-based smart systems and communication networks. Our objective is to balance accuracy, computational efficiency, and real-world deployability. To that end, we develop a lightweight 1D-CNN, optimized via hardware-aware neural architecture search (HW-NAS), which achieves 9 6. 5 9 % accuracy on the ISCX VPN-NonVPN dataset with only 88.26 K parameters, a 20.12 K maximum tensor size, and 10.08 M floating-point operations (FLOPs). Moreover, it generalizes across various TC tasks, with accuracies ranging from 94 % to 99 %. To enable deployment, the model is quantized to INT8, suffering only a marginal 1 - 2 % accuracy drop relative to its Float32 counterpart. We evaluate real-world inference performance on two microcontrollers: the high-performance STM32F746G-DISCO and the cost-sensitive Nucleo-F401RE. The deployed model achieves inference latencies of 31.43 ms and 115.40 ms, with energy consumption of 7.86 mJ and 29.10 mJ per inference, respectively. These results demonstrate the feasibility of on-device encrypted traffic analysis, paving the way for scalable, low-power IoT security solutions.

Energy-Efficient Deep Learning for Traffic Classification on Microcontrollers

Chehade, Adel;Ragusa, Edoardo;Gastaldo, Paolo;Zunino, Rodolfo
2025-01-01

Abstract

In this paper, we present a practical deep learning (DL) approach for energy-efficient traffic classification (TC) on resource-limited microcontrollers, which are widely used in IoT-based smart systems and communication networks. Our objective is to balance accuracy, computational efficiency, and real-world deployability. To that end, we develop a lightweight 1D-CNN, optimized via hardware-aware neural architecture search (HW-NAS), which achieves 9 6. 5 9 % accuracy on the ISCX VPN-NonVPN dataset with only 88.26 K parameters, a 20.12 K maximum tensor size, and 10.08 M floating-point operations (FLOPs). Moreover, it generalizes across various TC tasks, with accuracies ranging from 94 % to 99 %. To enable deployment, the model is quantized to INT8, suffering only a marginal 1 - 2 % accuracy drop relative to its Float32 counterpart. We evaluate real-world inference performance on two microcontrollers: the high-performance STM32F746G-DISCO and the cost-sensitive Nucleo-F401RE. The deployed model achieves inference latencies of 31.43 ms and 115.40 ms, with energy consumption of 7.86 mJ and 29.10 mJ per inference, respectively. These results demonstrate the feasibility of on-device encrypted traffic analysis, paving the way for scalable, low-power IoT security solutions.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11567/1320340
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 0
social impact