Data Engineering • 2026

DuckDB Hash Join Optimization

C++ database work on a cache-aligned Bloom filter and software prefetching inside DuckDB’s hash-join operator, evaluated across standard query workloads.

DuckDB — Hash join optimization. Conceptual illustration of data passing through a Bloom filter, prefetched cache blocks and a hash join.
AI-generated conceptual illustration.

Problem

Hash joins can waste memory bandwidth probing rows that will never match. An optimization needs to reduce that work without adding more overhead than it saves.

Approach

Implemented a cache-aligned Bloom filter and prefetching, then compared strategies across TPC-H, TPC-DS and IMDB. Benchmark results prompted a change from an adaptive-join approach to the Bloom filter design.

Impact

The recorded study reports 1.56× overall TPC-H speedup at 100 GB and covers 135 queries across three workloads. Results depend on the benchmark setup documented in the research write-up; the best selective-query result is separate from the overall figure.

Key Metrics

1.56×
Recorded overall TPC-H speedup · 100 GB
135 queries · 3 workloads
Recorded benchmark study

Technologies

C++DuckDBSQLTPC-HTPC-DS

Links

My Role

Worked on the database implementation, benchmark comparison and analysis documented in the research write-up.