Jesse Thomason, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, Raymond J. Mooney: Integrating Language and Vision to Generate Natural Language Descriptions of Videos in the Wild. COLING 2014: 1218-1227