DOI
10.34229/KCA2522-9664.26.5.3
UDC 004.932.72+004.855.5
V.M. Tereshchenko
Taras Shevchenko National University of Kyiv, Kyiv, Ukraine,
v.ter@knu.ua
P.V. Tytarchuk
Taras Shevchenko National University of Kyiv, Kyiv, Ukraine,
pavlo.tytarchuk@knu.ua
V.V. Hordiichuk
Taras Shevchenko National University of Kyiv, Kyiv, Ukraine,
vitaliyk-17@knu.ua
Y.V. Tereshchenko
Taras Shevchenko National University of Kyiv, Kyiv, Ukraine,
y.tereshchenko@knu.ua
SKETCH ANIMATION FROM TEXT: METHODS, CONTROL AND MOBILE DEPLOYMENT
Abstract. This article explores existing approaches and methods for generating animated sketch videos from text descriptions, focusing on the feasibility of on-device deployment and controlled generation mechanisms. The research aims to review and analyze current methods, identify critical challenges, and outline pathways toward fully autonomous on-device sketch animation systems. The key contributions include: a structured taxonomy of methods based on input type, control mechanism, and architecture; a critical analysis of their suitability for mobile deployment; the identification of key technical barriers and future research directions. The study analyzes publications from 2023–2025, covering diffusion-based approaches, controlled video generation frameworks, and mobile optimization techniques, including quantization, pruning, and knowledge distillation. The findings indicate that while current methods achieve high generation quality, they require significant computational resources that exceed mobile device capabilities by a factor of 10–100. Modern mobile neural processing units (NPUs) combined with model compression techniques offer promising paths for acceleration, yet substantial architectural innovations are still needed. Key research gaps include temporal consistency in lightweight models, trade-offs between fine-grained control and efficiency, and multi-object animation limitations.
Keywords: sketch animation, text-to-video generation, video diffusion models, mobile deployment, controllable generation, on-device inference, model compression, NPU.
full text
REFERENCES
- Soppari K., Nakka S.S.N., Mohammed A., Mogala M.K. A study on AI generated animated videos: An analytical review. World Journal of Advanced Research and Reviews. 2025. Vol. 26, N 2. P. 3347–3355. https://doi.org/10.30574/wjarr.2025.26.2.1954.
- OpenAI. Video generation models as world simulators. OpenAI Research. 2024. URL: https://openai.com/index/video-generation-models-as-world-simulators/.
- Google DeepMind. Veo: a text-to-video generation system. Technical Report. 2025. 7 p. URL: https://storage.googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.pdf.
- Kuaishou. Kling AI: Text-to-video generation model. 2024. URL: https://kling.kuaishou.com.
- Unlocking on-device generative AI with an NPU and heterogeneous computing. Technical White paper. San Diego, CA: Qualcomm Technologies, Inc., 2024. 18 p. URL: https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Unlocking-on-device-generative-AI-with-an-NPU-and-heterogeneous-computing.pdf.
- Tang M., Chen Y. AI and animated character design: efficiency, creativity, interactivity. The Frontiers of Society, Science and Technology. 2024. Vol. 6, N 1. P. 117–123. https://doi.org/10.25236/FSST.2024.060120.
- Blattmann A., Rombach R., Ling H., et al. Align your latents: High-resolution video synthesis with latent diffusion models. Proc. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (17–24 June 2023, Vancouver, BC, Canada). Vancouver, 2023. P. 22563–22575. https://doi.org/10.1109/CVPR52729.2023.02161.
- Peebles W., Xie S. Scalable diffusion models with transformers. Proc. 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (01–06 October 2023, Paris, France). Paris, 2023. P. 4172–4182. https://doi.org/10.1109/ICCV51070.2023.00387.
- Xing X., Wang C., Zhou H., et al. DiffSketcher: Text guided vector sketch synthesis through latent diffusion models. arXiv:2306.14685v5 [cs.CV] 8 Apr 2026. https://doi.org/10.48550/arXiv.2306.14685.
- Vinker Y., Pajouheshgar E., Bo J.Y., et al. CLIPasso: semantically-aware object sketching. ACM Transactions on Graphics. 2022. Vol. 41, Iss. 4. Article number 86. https://doi.org/10.1145/3528223.3530068.
- Gal R., Vinker Y., Alaluf Y., et al. Breathing life into sketches using text-to-video priors. Proc. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (16–22 June 2024, Seattle, WA, USA). Seattle, 2024. P. 4325–4336. https://doi.org/10.1109/CVPR52733.2024.00414.
- Bandyopadhyay H., Song Y.-Z. FlipSketch: Flipping static drawings to text-guided sketch animations. Proc. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (11–15 June 2025, Nashville, TN, USA). Nashville, 2025. P. 28394–28404. https://doi.org/10.1109/CVPR52734.2025.02644.
- Jiang L., Chen S., Wu B., et al. VidSketch: Hand-drawn sketch-driven video generation with diffusion control. Neural Networks. 2026. Vol. 196. Article number 108465. https://doi.org/10.1016/j.neunet.2025.108465.
- Meng Y., Ouyang H., Wang H., et al. AniDoc: Animation creation made easier. Proc. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (11–15 June 2025, Nashville, TN, USA). Nashville, 2025. P. 18187–18197. https://doi.org/10.1109/CVPR52734.2025.01695.
- Liu F.-L., Fu H., Wang X., et al. SketchVideo: Sketch-based video generation and editing. Proc. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (11–15 June 2025, Nashville, TN, USA). Nashville, 2025. P. 23379–23390. https://doi.org/10.1109/CVPR52734.2025.02177.
- Zhang L., Rao A., Agrawala M. Adding conditional control to text-to-image diffusion models. Proc. 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (02–06 October 2023, Paris, France). Paris, 2023. P. 3813–3824. https://doi.org/10.1109/ICCV51070.2023.00355.
- Peng B., Wang J., Zhang Y., et al. ControlNeXt: Powerful and efficient control for image and video generation. arXiv:2408.06070v3 [cs.CV] 8 Mar 2025. https://doi.org/10.48550/arXiv.2408.06070.
- Geng D., Herrmann C., Hur J., et al. Motion prompting: Controlling video generation with motion trajectories. Proc. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (11–15 June 2025, Nashville, TN, USA). Nashville, 2025. P. 1–12. https://doi.org/10.1109/CVPR52734.2025.00010.
- Chu R., He Y., Chen Z., et al. Wan-Move: Motion-controllable video generation via latent trajectory guidance. arXiv:2512.08765v1 [cs.CV] 9 Dec 2025. https://doi.org/10.48550/arXiv.2512.08765.
- Wang C., Gu J., Hu P., et al. EasyControl: Transfer ControlNet to video diffusion for controllable generation and interpolation. arXiv:2408.13005v2 [cs.CV] 16 Sep 2024. https://doi.org/10.48550/arXiv.2408.13005.
- Shi X., Huang Z., Wang F.-Y., et al. Motion-I2V: Consistent and controllable image-to-video generation with explicit motion modeling. Proc. SIGGRAPH’24: Special Interest Group on Computer Graphics and Interactive Techniques Conference (27 July – 1 August 2024, Denver, CO, USA). Denver, 2024. Article number 111. https://doi.org/10.1145/3641519.3657497.
- Musa A., Kakudi H.A., Hassan M., et al. Lightweight deep learning models for edge devices — a survey. International Journal of Computer Information Systems and Industrial Management Applications. 2025. Vol. 17. P. 189–206. https://doi.org/10.70917/ijcisim-2025-0014.
- Dantas P.V., Sabino da Silva W., Cordeiro L.C., et al. A comprehensive review of model compression techniques in machine learning. Applied Intelligence. 2024. Vol. 54. P. 11804–11844. https://doi.org/10.1007/s10489-024-05747-w.
- NVIDIA Model-Optimizer. GitHub Repository. 2025. URL: https://github.com/NVIDIA/TensorRT-Model-Optimizer.
- Lin J., Tang J., Tang H., et al. AWQ: Activation-aware weight quantization for LLM compression and acceleration. arXiv:2306.00978v6 [cs.CL] 25 Apr 2026. https://doi.org/10.48550/arXiv.2306.00978.
- Fang G., Ma X., Wang X. Structural pruning for diffusion models. Advances in Neural Information Processing Systems 36 (NeurIPS 2023) (10–16 December 2023, New Orleans, LA, USA). New Orleans, 2023. https://doi.org/10.48550/arXiv.2305.10924.
- Liu D., Zhu Y., Liu Z., et al. A survey of model compression techniques: Past, present, and future. Frontiers in Robotics and AI. 2025. Vol. 12. https://doi.org/10.3389/frobt.2025.1518965.
- Google Developers. TensorFlow Lite is now LiteRT. Google Developers Blog. 2024. URL: https://developers.googleblog.com/tensorflow-lite-is-now-litert/.
- Microsoft. ONNX Runtime: Cross-platform, high performance ML inferencing and training accelerator. 2025. URL: https://onnxruntime.ai/.
- Alibaba MNN: Mobile neural network framework. GitHub Repository. 2025. URL: https://github.com/alibaba/MNN.
- MLX: Efficient and flexible machine learning on Apple silicon. 2023. URL: https://github.com/ml-explore/mlx.
- Hong Y., Kim D. Performance and efficiency gains of NPU-based servers over GPUs for AI model inference. Systems. 2025. Vol. 13, Iss. 9. Article number 797. https://doi.org/10.3390/systems13090797.
- Xu D., Zhang H., Yang L., et al. Fast on-device LLM inference with NPUs. Proc. 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’25) (30 March – 3 April 2025, Rotterdam, Netherlands). Rotterdam, 2025. Vol. 1. P. 445–462. https://doi.org/10.1145/3669940.3707239.
- Liu J., Xin Z., Fu Y., et al. Multi-object sketch animation by scene decomposition and motion planning. arXiv:2503.19351v2 [cs.CV] 2 Aug 2025. https://doi.org/10.48550/arXiv.2503.19351.
- Yao Y., Yu T., Zhang A., et al. Efficient GPT-4V level multimodal large language model for deployment on edge devices. Nature Communications. 2025. Vol. 16. Article number 5509. https://doi.org/10.1038/s41467-025-61040-5.
- Lu Z., Li X., Cai D., et al. Demystifying small language models for edge deployment. Proc. 63rd Annual Meeting of the Association for Computational Linguistics (July 2025, Vienna, Austria). Vienna, 2025. Vol. 1: Long Papers. P. 14747–14764. https://doi.org/10.18653/v1/2025.acl-long.718.
- Liu X., Zhang W., Yan C., et al. Efficient real-time on-mobile video super-resolution with automatic evolutionary neural architecture search. In: Senn W. et al (Eds.). Artificial Neural Networks and Machine Learning — ICANN 2025. LNCS. Vol. 16069. P. 86–97. https://doi.org/10.1007/978-3-032-04546-1_8.