Word2Vec dictionary for 65000 Gutenberg E-books

dc.contributor.authorEgense, Thomas
dc.date.accessioned2018-06-26T11:04:29Z
dc.date.available2018-06-26T11:04:29Z
dc.description.abstractDescription: 55,000 e-books from Project Gutenberg (http://www.gutenberg.org/). About 35.000 books are english, but over 50 different languages are represented. The word2vec algorithm does a good job at seperating the different languages, so it is almost like it is 50 different word2vec dictionaries. Corpus size: 30GB of text spliteded into in 230 million sentences sentences with all punctuations removed. Word2Vec takes about 1.5 week/CPU time to build the dictionary. Word2Vec parameters: Software implementation:Google Model: Skip-Gram Word window size: 5 Iterations: 10 Minimum word frequency: 100 Dimensions:300 Ouput format:text Word2vec dictionary file: 1.4 million different words, from over 50 different languages. Requirements: Opening the dictionary in a word2vec implementation will require 16GB of memory.en_US
dc.identifier.urihttps://loar.kb.dk/handle/1902/327
dc.identifier.urihttp://dx.doi.org/10.21994/loar157
dc.language.isoenen_US
dc.rightsCC Public Domain*
dc.rights.urihttps://creativecommons.org/publicdomain/mark/1.0/deed.en*
dc.subjectword2vecen_US
dc.subjectmachine learningen_US
dc.subjectNLPen_US
dc.subjectGutenbergen_US
dc.titleWord2Vec dictionary for 65000 Gutenberg E-booksen_US
dc.typeDataseten_US
Files
Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
gutenberg_65K_books.txt
Size:
7.99 GB
Format:
Plain Text
Description:
Word2Vec dictionary file in text format for 65K Gutenberg E-books
License bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
4.47 KB
Format:
Item-specific license agreed upon to submission
Description:
Collections