Temporal Gradient Oscillation with Accuracy Recovery Mechanism for Efficient BERT-Based Text Classification
Abstract
Keywords
Full Text:
PDFReferences
L. Xue et al., “ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 291–306, Mar. 2022, doi: 10.1162/tacl_a_00461.
J. Yun, M. Kim, and Y. Kim, “Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore: Association for Computational Linguistics, 2023, pp. 13617–13628. doi: 10.18653/v1/2023.findings-emnlp.909.
Z. Zhan et al., “Exploring Token Pruning in Vision State Space Models”.
M. Kim, S. Gao, Y.-C. Hsu, Y. Shen, and H. Jin, “Token Fusion: Bridging the Gap between Token Pruning and Token Merging,” in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA: IEEE, Jan. 2024, pp. 1372–1381. doi: 10.1109/WACV57701.2024.00141.
E. Dotan, G. Jaschek, T. Pupko, and Y. Belinkov, “Effect of tokenization on transformers for biological sequences,” Bioinformatics, vol. 40, no. 4, p. btae196, Mar. 2024, doi: 10.1093/bioinformatics/btae196.
H. Wang, B. Dedhia, and N. K. Jha, “Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 16070–16079. doi: 10.1109/CVPR52733.2024.01521.
J. Zhu et al., “Exploring Dynamic Transformer for Efficient Object Tracking,” IEEE Trans. Neural Netw. Learning Syst., vol. 36, no. 8, pp. 15502–15514, Aug. 2025, doi: 10.1109/TNNLS.2025.3545752.
W. Li et al., “TokenPacker: Efficient Visual Projector for Multimodal LLM,” International Journal of Computer Vision, vol. 133, no. 10, pp. 6794–6812, Oct. 2025, doi: 10.1007/s11263-025-02491-7.
X. Liu, T. Wu, and G. Guo, “Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, Macau, SAR China: International Joint Conferences on Artificial Intelligence Organization, Aug. 2023, pp. 1222–1230. doi: 10.24963/ijcai.2023/136.
K. Xu et al., “LPViT: Low-Power Semi-structured Pruning for Vision Transformers,” Apr. 15, 2025, arXiv: arXiv:2407.02068. doi: 10.48550/arXiv.2407.02068.
S. Wei, T. Ye, S. Zhang, Y. Tang, and J. Liang, “Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 2092–2101. doi: 10.1109/CVPR52729.2023.00208.
H. C. Tran et al., “Accelerating Transformers with Spectrum-Preserving Token Merging,” in Advances in Neural Information Processing Systems, 2024. doi: 10.52202/079017-0969.
S. Kim et al., “Learned Token Pruning for Transformers,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Washington DC USA: ACM, Aug. 2022, pp. 784–794. doi: 10.1145/3534678.3539260.
J. Li et al., “Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference,” in KDD ’23: The 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, ACM, 2023, pp. 1280–1290. doi: 10.1145/3580305.3599284.
M. Guo et al., “LongT5: Efficient Text-To-Text Transformer for Long Sequences,” in Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, United States: Association for Computational Linguistics, 2022, pp. 724–736. doi: 10.18653/v1/2022.findings-naacl.55.
J. Yun, M. Kim, and Y. Kim, “Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computational Linguistics, 2023, pp. 13617–13628. doi: 10.18653/v1/2023.findings-emnlp.909.
E. Dotan, G. Jaschek, T. Pupko, and Y. Belinkov, “Effect of tokenization on transformers for biological sequences,” Bioinformatics, vol. 40, no. 4, p. btae196, 2024, doi: 10.1093/bioinformatics/btae196.
L. Wei, W. Yimin, and Z. Hao, “Accelerating NLP with Token Pruning: A Survey of Methods and Applications,” TechRxiv, 2025, doi: 10.36227/techrxiv.174195729.93551382/v1.
M. Guo et al., “LongT5: Efficient Text-To-Text Transformer for Long Sequences,” in Findings of the Association for Computational Linguistics: NAACL 2022, Association for Computational Linguistics, 2022, pp. 724–736. doi: 10.18653/v1/2022.findings-naacl.55.
N. Zafarmomen and V. Samadi, “Can large language models effectively reason about adverse weather conditions?,” Environmental Modelling and Software, vol. 188, no. January, p. 106421, 2025, doi: 10.1016/j.envsoft.2025.106421.
S. Lee et al., “Entity-enhanced BERT for medical specialty prediction based on clinical questionnaire data,” PLoS ONE, vol. 20, no. 1 January, pp. 1–18, 2025, doi: 10.1371/journal.pone.0317795.
S. Kim et al., “Learned Token Pruning for Transformers,” in KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, ACM, 2022, pp. 784–794. doi: 10.1145/3534678.3539260.
C. Lee, F. F. Khan, R. B. Brufau, K. Ding, and V. Narayanan, “Token and Head Adaptive Transformers for Efficient Natural Language Processing”.
X. Jie, Y. Yang, and Y. Jianhong, “Advancing Transformer Efficiency with Token Pruning,” Preprints, 2025, doi: 10.20944/preprints202503.1577.v1.
Y. Yan, “Attribution-Driven Adaptive Token Pruning for Transformers,” no. NeurIPS, pp. 1–22, 2025.
S. B. Belhaouari and I. Kraidia, “Efficient self-attention with smart pruning for sustainable large language models,” Scientific Reports, vol. 15, no. 1, p. 10171, 2025, doi: 10.1038/s41598-025-92586-5.
W. Lei, Z. Wei, and S. Xiuying, “Reducing Computational Overhead in Transformers with Token Pruning,” 2025, SSRN. doi: 10.2139/ssrn.5171805.
Z. Wen, Y. Gao, W. Li, C. He, and L. Zhang, “Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?”.
H. Wang, B. Dedhia, and N. K. Jha, “Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2024, pp. 16070–16079. doi: 10.1109/CVPR52733.2024.01521.
Z. Zhan et al., “Exploring Token Pruning in Vision State Space Models”.
DOI: https://doi.org/10.30743/jet.v11i2.13472
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Indra Listiawan

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





