Dekun Xie
GTZAN [4] is used as datasets for music genre classfication; the models are trained with AlexNet [1] and ResNet [2].
The datasets used in project as well as the trained models can be download from: https://1drv.ms/u/s!AgDWRokUsfkDh_5r0hOjoWZcHvgSSQ?e=Ogw4EF
[1] Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
[2] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
[3] Dong, M. (2018). Convolutional Neural Network Achieves Human-level Accuracy in Music Genre Classification. http://arxiv.org/abs/1802.09697
[4] A. Olteanu. Gtzan dataset - music genre classification. [Online]. Available: https://www.kaggle.com/andradaolteanu/gtzan-dataset-musicgenre-classification
[5] Zhang, A., Lipton, Z. C., Li, M., & Smola, A. J. (2021). Dive into deep learning. arXiv preprint arXiv:2106.11342.
[6] Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
[7] https://en.wikipedia.org/wiki/AlexNet
[8] https://en.wikipedia.org/wiki/Residual_neural_network
[9] Santurkar, S., Tsipras, D., Ilyas, A., & Mit, A. M. ? A. (n.d.). How Does Batch Normalization Help Optimization?
[10] George Tzanetakis and Perry Cook. Musical genre classification of audio signals.IEEE Transactions onspeech and audio processing, 10(5):293?302, 2002