Word2Vec dictionary for 65000 Gutenberg E-books

Egense, Thomas

Word2Vec dictionary for 65000 Gutenberg E-books

Files

gutenberg_65K_books.txt(7.99 GB)

Authors

Egense, Thomas

Abstract

Description: 55,000 e-books from Project Gutenberg (http://www.gutenberg.org/). About 35.000 books are english, but over 50 different languages are represented. The word2vec algorithm does a good job at seperating the different languages, so it is almost like it is 50 different word2vec dictionaries. Corpus size: 30GB of text spliteded into in 230 million sentences sentences with all punctuations removed. Word2Vec takes about 1.5 week/CPU time to build the dictionary. Word2Vec parameters: Software implementation:Google Model: Skip-Gram Word window size: 5 Iterations: 10 Minimum word frequency: 100 Dimensions:300 Ouput format:text Word2vec dictionary file: 1.4 million different words, from over 50 different languages. Requirements: Opening the dictionary in a word2vec implementation will require 16GB of memory.

Keywords

word2vec, machine learning, NLP, Gutenberg

URI

https://loar.kb.dk/handle/1902/327
http://dx.doi.org/10.21994/loar157

Collections

Machine learning

Full item page