A (2+1)D Attention Convolutional Neural Network for Video Prediction
Tác giả
Tóm tắt
Tài liệu tham khảo
[1] Aigner S, Körner M (2018) FutureGAN: anticipating the future frames of video sequences using spatio-temporal 3d convolutions in progressively growing GANs. arXiv preprint arXiv:1810.01325
[2] Byeon W, Wang Q, Kumar Srivastava R, Koumoutsakos P (2018) ContextVP: fully context-aware video prediction. In: ECCV, pp 753–769
[3] Castrejon L, Ballas N, Courville A (2019) Improved conditional VRNNs for video prediction. In: ICCV, pp 7608–7617
[4] Diba A, Fayyaz M, Sharma V, Mahdi Arzani M, Yousefzadeh R, Gall J, Van Gool L (2018) Spatio-temporal channel correlation networks for action classification. In: ECCV, pp 284–299
[5] Elafi I, Jedra M, Zahid N (2016) Unsupervised detection and tracking of moving objects for video surveillance applications. Pattern Recogn Lett 84:70–77
[6] Gao H, Xu H, Cai QZ, Wang R, Yu F, Darrell T (2019) Disentangling propagation and generation for video prediction. In: ICCV, pp 9006–9015
[7] Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. In: Advances in neural information processing systems, vol 27
[8] Hao Z, Huang X, Belongie S (2018) Controllable video generation with sparse trajectories. In: CVPR, pp 7854–7863
[9] Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: CVPR, June 2018
[10] Jin B, Hu Y, Tang Q, Niu J, Shi Z, Han Y, Li X (2020) Exploring spatial-temporal multi-frequency analysis for high-fidelity and temporal-consistency video prediction. In: CVPR, pp 4554–4563
[11] Johnson J, Alahi A, Fei-Fei L (2016) Perceptual losses for real-time style transfer and super-resolution. In: ECCV. Springer, Berlin, pp 694–711
[12] Kwon YH, Park MG (2019) Predicting future frames using retrospective cycle GAN. In: CVPR, pp 1811–1820
[13] Lee AX, Zhang R, Ebert F, Abbeel P, Finn C, Levine S (2018) Stochastic adversarial video prediction. arXiv preprint arXiv:1804.01523
[14] Lee W, Jung W, Zhang H, Chen T, Koh JY, Huang T, Yoon H, Lee H, Hong S (2021) Revisiting hierarchical approach for persistent long-term video prediction. arXiv preprint arXiv:2104.06697
[15] Liang X, Lee L, Dai W, Xing EP (2017) Dual motion GAN for future-flow embedded video prediction. In: ICCV, pp 1744–1752
[16] Liu B, Chen Y, Liu S, Kim HS (2021) Deep learning in latent space for video prediction and compression. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 701–710
[17] Lotter W, Kreiman G, Cox D (2016) Deep predictive coding networks for video prediction and unsupervised learning. arXiv preprint arXiv:1605.08104
[18] Mathieu M, Couprie C, LeCun Y (2015) Deep multi-scale video prediction beyond mean square error. arXiv preprint arXiv:1511.05440
[19] Oprea S, Martinez-Gonzalez P, Garcia-Garcia A, Castro-Vargas JA, Orts-Escolano S, Garcia-Rodriguez J, Argyros A (2020) A review on deep learning techniques for video prediction. arXiv preprint arXiv:2004.05214
[20] Rasouli A (2020) Deep learning for vision-based prediction: a survey. arXiv preprint arXiv:2007.00095
[21] Soomro K, Zamir AR, Shah M (2021) A dataset of 101 human action classes from videos in the wild. In: Center for research in computer vision, vol 2
[22] Srivastava N, Mansimov E, Salakhudinov R (2015) Unsupervised learning of video representations using LSTMs. In: ICML, pp 843–852
[23] Straka Z, Svoboda T, Hoffmann M (2020) PreCNet: next frame video prediction based on predictive coding. arXiv preprint arXiv:2004.14878
[24] Tran D, Wang H, Torresani L, Ray J, LeCun Y, Paluri M (2018) A closer look at spatiotemporal convolutions for action recognition. In: CVPR, pp 6450–6459
[25] Wang Y, Jiang L, Yang MH, Li LJ, Long M, Fei-Fei L (2018) Eidetic 3d LSTM: a model for video prediction and beyond. In: ICLR
[26] Woo S, Park J, Lee JY, So Kweon I (2018) CBAM: convolutional block attention module. In: ECCV, pp 3–19
[27] Zeng KH, Shen WB, Huang DA, Sun M, Carlos Niebles J (2017) Visual forecasting by imitating dynamics in natural sequences. In: ICCV, pp 2999–3008
[28] Zhang C, Chen T, Liu H, Shen Q, Ma Z (2019) Looking-ahead: neural future video frame prediction. In: ICIP. IEEE, pp 1975–1979
[29] Zhang R, Isola P, Efros AA, Shechtman E, Wang O (2018) The unreasonable effectiveness of deep features as a perceptual metric. In: CVPR, pp 586–595