200
Bibliography
21. Mohri, M., Rostamizadeh, A., Talwalkar, A.: Foundations of Machine Learning. The MIT
Press (2018)
22. Akaike, H.: Information theory and an extension of the maximum likelihood principle. In:
Selected Papers of Hirotugu Akaike, pp. 199–213. Springer (1998)
23. Borsanyi, S., et al.: Ab initio calculation of the neutron-proton mass difference. Science 347,
1452–1455 (2015)
24. He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In:
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–
778 (2016)
25. Shalev-Shwartz, S., Ben-David, S.: Understanding Machine Learning: From Theory to
Algorithms. Cambridge University Press (2014)
26. Deng, J., Dong, W., Socher, R., Li, L., Li, K., Fei-Fei, L.: Imagenet: a large-scale hierarchical
image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp.
248–255 (2009)
27. Kawaguchi, K., Kaelbling, L.P., Bengio, Y.: Generalization in deep learning. arXiv preprint
arXiv:1710.05468 (2017)
28. Goodfellow, I., Bengio, Y., Courville, A.: Deep Learning. MIT Press (2016)
29. Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. In:
Proceedings of the 27th International Conference on Machine Learning (ICML-10), pp. 807–
814 (2010)
30. Teh, Y.W., Hinton, G.E.: Rate-coded restricted boltzmann machines for face recognition. In:
Advances in Neural Information Processing Systems, pp. 908–914 (2001)
31. Rumelhart, D.E., Hinton, G.E., Williams, R.J.: Learning internal representations by error
propagation. Technical report, California Univ San Diego La Jolla Inst for Cognitive Science
(1985)
32. Cybenko, G.: Approximation by superpositions of a sigmoidal function. Math. Control
Signals Syst. 2(4), 303–314 (1989)
33. Nielsen, M.: A visual proof that neural nets can compute any function. http://
neuralnetworksanddeeplearning.com/chap4.html.
34. Lee, H., Ge, R., Ma, T., Risteski, A., Arora, S.: On the ability of neural nets to express
distributions. arXiv preprint arXiv:1702.07028 (2017)
35. Sonoda, S., Murata, N.: Neural network with unbounded activation functions is universal
approximator. Appl. Comput. Harmon Anal. 43(2), 233–268 (2017)
36. LeCun, Y., Bottou, L., Bengio, Y., Haffner, P., et al.: Gradient-based learning applied to
document recognition. Proc. IEEE 86(11), 2278–2324 (1998)
37. Hubel, D.H., Wiesel, T.N.: Receptive fields, binocular interaction and functional architecture
in the cat’s visual cortex. J. Physiol. 160(1), 106–154 (1962)
38. Fukushima, K., Miyake, S.: Neocognitron: a new algorithm for pattern recognition tolerant of
deformations and shifts in position. Pattern Recognit. 15(6), 455–469 (1982)
39. Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional
neural networks. In: Advances in Neural Information Processing Systems, pp. 1097–1105
(2012)
40. Sabour, S., Frosst, N., Hinton, G.E.: Dynamic routing between capsules. In: Advances in
Neural Information Processing Systems, pp. 3856–3866 (2017)
41. Dumoulin, V., Visin, F.: A guide to convolution arithmetic for deep learning. arXiv preprint
arXiv:1603.07285 (2016)
42. Radford, A., Metz, L., Chintala, S.: Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434 (2015)
43. Siegelmann, H.T., Sontag, E.D.: On the computational power of neural nets. J. Comput. Syst.
Sci. 50(1), 132–150 (1995)
44. Hochreiter, S.: Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut
für Informatik (1991)
45. Hochreiter, S., Bengio, Y., Frasconi, P., Schmidhuber, J., et al.: Gradient flow in recurrent nets:
the difficulty of learning long-term dependencies. In: A Field Guide to Dynamical Recurrent
Bibliography
21. Mohri, M., Rostamizadeh, A., Talwalkar, A.: Foundations of Machine Learning. The MIT
Press (2018)
22. Akaike, H.: Information theory and an extension of the maximum likelihood principle. In:
Selected Papers of Hirotugu Akaike, pp. 199–213. Springer (1998)
23. Borsanyi, S., et al.: Ab initio calculation of the neutron-proton mass difference. Science 347,
1452–1455 (2015)
24. He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In:
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–
778 (2016)
25. Shalev-Shwartz, S., Ben-David, S.: Understanding Machine Learning: From Theory to
Algorithms. Cambridge University Press (2014)
26. Deng, J., Dong, W., Socher, R., Li, L., Li, K., Fei-Fei, L.: Imagenet: a large-scale hierarchical
image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp.
248–255 (2009)
27. Kawaguchi, K., Kaelbling, L.P., Bengio, Y.: Generalization in deep learning. arXiv preprint
arXiv:1710.05468 (2017)
28. Goodfellow, I., Bengio, Y., Courville, A.: Deep Learning. MIT Press (2016)
29. Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. In:
Proceedings of the 27th International Conference on Machine Learning (ICML-10), pp. 807–
814 (2010)
30. Teh, Y.W., Hinton, G.E.: Rate-coded restricted boltzmann machines for face recognition. In:
Advances in Neural Information Processing Systems, pp. 908–914 (2001)
31. Rumelhart, D.E., Hinton, G.E., Williams, R.J.: Learning internal representations by error
propagation. Technical report, California Univ San Diego La Jolla Inst for Cognitive Science
(1985)
32. Cybenko, G.: Approximation by superpositions of a sigmoidal function. Math. Control
Signals Syst. 2(4), 303–314 (1989)
33. Nielsen, M.: A visual proof that neural nets can compute any function. http://
neuralnetworksanddeeplearning.com/chap4.html.
34. Lee, H., Ge, R., Ma, T., Risteski, A., Arora, S.: On the ability of neural nets to express
distributions. arXiv preprint arXiv:1702.07028 (2017)
35. Sonoda, S., Murata, N.: Neural network with unbounded activation functions is universal
approximator. Appl. Comput. Harmon Anal. 43(2), 233–268 (2017)
36. LeCun, Y., Bottou, L., Bengio, Y., Haffner, P., et al.: Gradient-based learning applied to
document recognition. Proc. IEEE 86(11), 2278–2324 (1998)
37. Hubel, D.H., Wiesel, T.N.: Receptive fields, binocular interaction and functional architecture
in the cat’s visual cortex. J. Physiol. 160(1), 106–154 (1962)
38. Fukushima, K., Miyake, S.: Neocognitron: a new algorithm for pattern recognition tolerant of
deformations and shifts in position. Pattern Recognit. 15(6), 455–469 (1982)
39. Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional
neural networks. In: Advances in Neural Information Processing Systems, pp. 1097–1105
(2012)
40. Sabour, S., Frosst, N., Hinton, G.E.: Dynamic routing between capsules. In: Advances in
Neural Information Processing Systems, pp. 3856–3866 (2017)
41. Dumoulin, V., Visin, F.: A guide to convolution arithmetic for deep learning. arXiv preprint
arXiv:1603.07285 (2016)
42. Radford, A., Metz, L., Chintala, S.: Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434 (2015)
43. Siegelmann, H.T., Sontag, E.D.: On the computational power of neural nets. J. Comput. Syst.
Sci. 50(1), 132–150 (1995)
44. Hochreiter, S.: Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut
für Informatik (1991)
45. Hochreiter, S., Bengio, Y., Frasconi, P., Schmidhuber, J., et al.: Gradient flow in recurrent nets:
the difficulty of learning long-term dependencies. In: A Field Guide to Dynamical Recurrent
