mirror of
https://github.com/google/cdc-file-transfer.git
synced 2026-09-13 01:10:44 +03:00
Change fastcdc to a better and simpler algorithm. (#79)
This CL changes the chunking algorithm from "normalized chunking" to simple "regression chunking", and changes the has criteria from 'hash&mask' to 'hash<=threshold'. These are all ideas taken from testing and analysis done at https://github.com/dbaarda/rollsum-chunking/blob/master/RESULTS.rst Regression chunking was introduced in https://www.usenix.org/system/files/conference/atc12/atc12-final293.pdf The algorithm uses an arbitrary number of regressions using power-of-2 regression target lengths. This means we can use a simple bitmask for the regression hash criteria. Regression chunking yields high deduplication rates even for lower max chunk sizes, so that the cdc_stream max chunk can be reduced to 512K from 1024K. This fixes potential latency spikes from large chunks.
This commit is contained in:
@@ -140,8 +140,7 @@ Indexer::Impl::Impl(const IndexerConfig& cfg,
|
||||
fastcdc::Config ccfg(cfg_.min_chunk_size, cfg_.avg_chunk_size,
|
||||
cfg_.max_chunk_size);
|
||||
Indexer::Chunker chunker(ccfg, nullptr);
|
||||
cfg_.mask_s = chunker.Stage(0).mask;
|
||||
cfg_.mask_l = chunker.Stage(chunker.StagesCount() - 1).mask;
|
||||
cfg_.threshold = chunker.Threshold();
|
||||
// Collect inputs.
|
||||
for (auto it = inputs.begin(); it != inputs.end(); ++it) {
|
||||
inputs_.push(*it);
|
||||
@@ -368,8 +367,7 @@ IndexerConfig::IndexerConfig()
|
||||
max_chunk_size(0),
|
||||
max_chunk_size_step(0),
|
||||
num_threads(0),
|
||||
mask_s(0),
|
||||
mask_l(0) {}
|
||||
threshold(0) {}
|
||||
|
||||
Indexer::Indexer() : impl_(nullptr) {}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user