Temporal Gradient Oscillation with Accuracy Recovery Mechanism for Efficient BERT-Based Text Classification

Indra Listiawan, Ema Utami, Ema Utami, Kusrini Kusrini, Kusrini Kusrini, Arief Setyanto, Arief Setyanto


Abstract


Large Language Models (LLMs) such as BERT have demonstrated impressive performance across various NLP tasks, yet their high computational cost poses challenges for deployment in resource-constrained environments. This paper proposes a dynamic temporal token pruning approach based on gradient oscillation monitoring, where token importance is estimated from the temporal variability of gradient signals during training. Tokens exhibiting low gradient oscillation are selectively pruned to reduce effective input length. To mitigate potential performance degradation caused by aggressive pruning, an accuracy recovery mechanism based on lightweight re-finetuning is introduced. Experimental results on benchmark sentiment classification datasets, including IMDB and SST-2, demonstrate that the proposed method substantially reduces the number of input tokens while maintaining or recovering predictive performance. These results indicate that gradient oscillation provides a viable signal for token-level efficiency, achieving a favorable trade-off between input sparsity and model accuracy without modifying the underlying Transformer architecture.

Keywords


Token Pruning; Osilasi Gradient; Dynamic Temporal Token; BERT

Full Text:

PDF

References


L. Xue et al., “ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 291–306, Mar. 2022, doi: 10.1162/tacl_a_00461.

J. Yun, M. Kim, and Y. Kim, “Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore: Association for Computational Linguistics, 2023, pp. 13617–13628. doi: 10.18653/v1/2023.findings-emnlp.909.

Z. Zhan et al., “Exploring Token Pruning in Vision State Space Models”.

M. Kim, S. Gao, Y.-C. Hsu, Y. Shen, and H. Jin, “Token Fusion: Bridging the Gap between Token Pruning and Token Merging,” in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA: IEEE, Jan. 2024, pp. 1372–1381. doi: 10.1109/WACV57701.2024.00141.

E. Dotan, G. Jaschek, T. Pupko, and Y. Belinkov, “Effect of tokenization on transformers for biological sequences,” Bioinformatics, vol. 40, no. 4, p. btae196, Mar. 2024, doi: 10.1093/bioinformatics/btae196.

H. Wang, B. Dedhia, and N. K. Jha, “Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 16070–16079. doi: 10.1109/CVPR52733.2024.01521.

J. Zhu et al., “Exploring Dynamic Transformer for Efficient Object Tracking,” IEEE Trans. Neural Netw. Learning Syst., vol. 36, no. 8, pp. 15502–15514, Aug. 2025, doi: 10.1109/TNNLS.2025.3545752.

W. Li et al., “TokenPacker: Efficient Visual Projector for Multimodal LLM,” International Journal of Computer Vision, vol. 133, no. 10, pp. 6794–6812, Oct. 2025, doi: 10.1007/s11263-025-02491-7.

X. Liu, T. Wu, and G. Guo, “Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, Macau, SAR China: International Joint Conferences on Artificial Intelligence Organization, Aug. 2023, pp. 1222–1230. doi: 10.24963/ijcai.2023/136.

K. Xu et al., “LPViT: Low-Power Semi-structured Pruning for Vision Transformers,” Apr. 15, 2025, arXiv: arXiv:2407.02068. doi: 10.48550/arXiv.2407.02068.

S. Wei, T. Ye, S. Zhang, Y. Tang, and J. Liang, “Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 2092–2101. doi: 10.1109/CVPR52729.2023.00208.

H. C. Tran et al., “Accelerating Transformers with Spectrum-Preserving Token Merging,” in Advances in Neural Information Processing Systems, 2024. doi: 10.52202/079017-0969.

S. Kim et al., “Learned Token Pruning for Transformers,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Washington DC USA: ACM, Aug. 2022, pp. 784–794. doi: 10.1145/3534678.3539260.

J. Li et al., “Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference,” in KDD ’23: The 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, ACM, 2023, pp. 1280–1290. doi: 10.1145/3580305.3599284.

M. Guo et al., “LongT5: Efficient Text-To-Text Transformer for Long Sequences,” in Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, United States: Association for Computational Linguistics, 2022, pp. 724–736. doi: 10.18653/v1/2022.findings-naacl.55.

J. Yun, M. Kim, and Y. Kim, “Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computational Linguistics, 2023, pp. 13617–13628. doi: 10.18653/v1/2023.findings-emnlp.909.

E. Dotan, G. Jaschek, T. Pupko, and Y. Belinkov, “Effect of tokenization on transformers for biological sequences,” Bioinformatics, vol. 40, no. 4, p. btae196, 2024, doi: 10.1093/bioinformatics/btae196.

L. Wei, W. Yimin, and Z. Hao, “Accelerating NLP with Token Pruning: A Survey of Methods and Applications,” TechRxiv, 2025, doi: 10.36227/techrxiv.174195729.93551382/v1.

M. Guo et al., “LongT5: Efficient Text-To-Text Transformer for Long Sequences,” in Findings of the Association for Computational Linguistics: NAACL 2022, Association for Computational Linguistics, 2022, pp. 724–736. doi: 10.18653/v1/2022.findings-naacl.55.

N. Zafarmomen and V. Samadi, “Can large language models effectively reason about adverse weather conditions?,” Environmental Modelling and Software, vol. 188, no. January, p. 106421, 2025, doi: 10.1016/j.envsoft.2025.106421.

S. Lee et al., “Entity-enhanced BERT for medical specialty prediction based on clinical questionnaire data,” PLoS ONE, vol. 20, no. 1 January, pp. 1–18, 2025, doi: 10.1371/journal.pone.0317795.

S. Kim et al., “Learned Token Pruning for Transformers,” in KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, ACM, 2022, pp. 784–794. doi: 10.1145/3534678.3539260.

C. Lee, F. F. Khan, R. B. Brufau, K. Ding, and V. Narayanan, “Token and Head Adaptive Transformers for Efficient Natural Language Processing”.

X. Jie, Y. Yang, and Y. Jianhong, “Advancing Transformer Efficiency with Token Pruning,” Preprints, 2025, doi: 10.20944/preprints202503.1577.v1.

Y. Yan, “Attribution-Driven Adaptive Token Pruning for Transformers,” no. NeurIPS, pp. 1–22, 2025.

S. B. Belhaouari and I. Kraidia, “Efficient self-attention with smart pruning for sustainable large language models,” Scientific Reports, vol. 15, no. 1, p. 10171, 2025, doi: 10.1038/s41598-025-92586-5.

W. Lei, Z. Wei, and S. Xiuying, “Reducing Computational Overhead in Transformers with Token Pruning,” 2025, SSRN. doi: 10.2139/ssrn.5171805.

Z. Wen, Y. Gao, W. Li, C. He, and L. Zhang, “Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?”.

H. Wang, B. Dedhia, and N. K. Jha, “Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2024, pp. 16070–16079. doi: 10.1109/CVPR52733.2024.01521.

Z. Zhan et al., “Exploring Token Pruning in Vision State Space Models”.




DOI: https://doi.org/10.30743/jet.v11i2.13472

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Indra Listiawan

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.