Commit Graph

126 Commits

Author SHA1 Message Date
ce8e56ee18 Rewrite the indexer to use one MTBL by database
This allows us to avoid prefixing keys and appending into LMDB databases
2020-10-04 17:04:33 +02:00
c4b0c57059 Reduce the default indexer max-memory parameter 2020-10-02 16:47:41 +02:00
f6a8096720 Rename the quartile as percentiles 25th, 50th and 75th 2020-10-02 16:46:07 +02:00
891e0188dd Introduce the database-stats infos subcommand 2020-10-02 16:46:07 +02:00
079742b4d3 Clean up the stats and size of database infos subcommands 2020-10-02 16:46:06 +02:00
d0c73564b1 Use the CboRoaringBitmapCodec for the word pair proximity docids 2020-10-02 16:46:06 +02:00
4eda149ffa Rename the BoRoaringBitmap codec 2020-10-02 16:46:06 +02:00
ac84db2506 Move the words pairs proximities average into the stats infos subcommand 2020-10-02 16:46:06 +02:00
30755e31e7 Introduce the words pairs proximities stats info subcommand 2020-10-02 16:46:06 +02:00
bc35c9a598 Introduce the size_of_database infos subcommand 2020-10-02 16:46:05 +02:00
58237bd67f Introduce the average-number-of-document-by-word-pair-proximity infos subcommand 2020-09-29 18:32:48 +02:00
991be8950e Rename the subcommand into average-number-of-positions-by-word-by-doc 2020-09-29 18:15:44 +02:00
54370e228a Search for documents with longer proximities until we find enough 2020-09-29 17:37:14 +02:00
68f4af7d2e Improve the display of the number of processed documents 2020-09-29 16:08:58 +02:00
59a127d022 Improve the indexing process
We now store the words pairs proximity in a cache and only compute the
shortest proximity between pairs of words in a document.
2020-09-29 15:09:18 +02:00
d8354f6f02 Fix the word_docids capacity limit detection 2020-09-27 11:52:05 +02:00
25b2853b70 Move the words pairs proximities compute into the write document function 2020-09-23 15:02:40 +02:00
ed05999f63 Replace the arc cache by a simple linked hash map 2020-09-23 14:50:52 +02:00
4d22d80281 Display only the key on heed error 2020-09-23 14:13:51 +02:00
b597a92487 Add a default max-memory value to the indexer 2020-09-23 12:00:36 +02:00
31224a8425 Index the word pair proximities for both orders of the pair 2020-09-22 14:49:22 +02:00
a58ae5eb2a Introduce the word-pair-proximities-docids infos subcommand 2020-09-22 14:04:34 +02:00
d6fa9c0414 Index the intra documents word pair proximities 2020-09-22 14:04:33 +02:00
15208c7d3d Simplify the indexer record loop 2020-09-22 10:33:30 +02:00
e5adfaade0 Replace the token filter by a filter mapper 2020-09-22 10:24:31 +02:00
d21c80b865 Apply the chunk compression parameters on all the MTBL writers 2020-09-21 18:30:54 +02:00
944df52e2a Simplify the indexer main loop 2020-09-21 14:59:48 +02:00
d5e5baa20f Bump the oxidized-mtbl dependency 2020-09-10 13:29:12 +02:00
ad11c5fb3f Introduce the words-docids command for the infos binary 2020-09-07 22:36:35 +02:00
5664c37539 Introduce an heed codec that reduce the size of small amount of serialized integers 2020-09-07 20:06:23 +02:00
3e2250423c Introduce the average-number-of-positions infos subcommand 2020-09-07 15:26:42 +02:00
ea605b499c Introduce two new infos subcommands 2020-09-07 14:56:48 +02:00
bb1ab428db Use another function to define the proximity 2020-09-06 17:55:07 +02:00
dec460ce52 Fix the infos binary and add commands 2020-09-06 17:14:20 +02:00
daa3673c1c Invert the word docid positions key order 2020-09-06 10:30:53 +02:00
c2405bcae2 Prefer using the word_docids db to create the words-fst 2020-09-06 10:23:56 +02:00
4ca9472e02 Fix the minimum proximity len 2020-09-06 10:19:34 +02:00
dc88a86259 Store the word positions under the documents 2020-09-05 18:03:06 +02:00
580ed1119a Make the engine to return csv string records as documents and headers 2020-08-31 19:02:00 +02:00
bad0663138 Come back to the old tokenizer 2020-08-31 13:34:38 +02:00
4afc4d0751 Use the groups of four positions to speed up disjunctions tests 2020-08-30 16:25:11 +02:00
605f75b56f Add the words grouped by four positions in the infos binary 2020-08-29 18:23:33 +02:00
ad5cafbfed Introduce a database to store docids in groups of four positions 2020-08-29 17:42:55 +02:00
3db517548d Move the documents back into the LMDB database 2020-08-29 15:14:04 +02:00
3fe497e129 Improve the Mtbl heed codec to only encode MTBL databases 2020-08-29 11:20:39 +02:00
21aafd603c Make sure the first document is associated to the document id 0 2020-08-29 10:56:40 +02:00
0a44ff86ab Put the documents MTBL back into LMDB
We makes sure to write the documents into a file before
memory mapping it and putting it into LMDB, this way we avoid
moving it to RAM
2020-08-28 15:43:24 +02:00
7cde312f14 Introduce the StrBEU32Codec heed codec 2020-08-28 14:16:37 +02:00
ba2eb0d7ad Take the words-fst into account when retrieving the biggests values 2020-08-26 14:36:22 +02:00
32da07ccee Introduce the word-positions-doc-ids and words-positions infos commands 2020-08-23 10:52:47 +02:00