TL;DR
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Polars has released version 2.0, making its streaming engine the default for collecting lazy queries and enabling initial spill-to-disk support. The release also promotes SQL support and adds a Map dtype; performance comparisons published by Polars are vendor-run benchmarks, not independent results.
Polars has released version 2.0, changing lazy query collection to use its streaming engine by default and enabling initial spill-to-disk support. The data-processing library says the release also makes SQL a first-class interface, adds a Map data type and introduces stricter expectations around data types and explicitness.
Under the new default, calling collect on a LazyFrame uses the streaming engine. Polars says that can reduce memory use and improve performance on many queries. However, the engine does not preserve row order by default for some operations, including joins, group-by operations and unpivoting. Users who need observable order for supported operations can request it with maintain_order=True.
Version 2.0 also enables out-of-core processing, which lets supported operations spill data to disk when memory use grows. According to Polars, spilling starts at about 80% of RAM, a threshold the project says may need tuning; the default disk budget is 64 GB. Current support includes sort, window functions and many expressions. Polars says out-of-core support for joins and group-by operations is planned, but is not available in this release.
Other changes include a native Map dtype for Arrow MapType data, with dictionary-like key and value operations. Polars describes SQL as a first-class part of the product and cites optimizer and engine work, including join reordering, common-subplan elimination and dynamic predicates or bloom filters. It also says stricter handling of data types and explicitness is intended to provide faster feedback during development.
Streaming Changes Query Defaults
The main practical change is not just a new query feature: streaming is now the default execution path when collecting a lazy query. That may affect both resource use and results where row order matters. Teams upgrading existing applications should check whether they rely on the previous ordering behavior for joins, group-by operations or unpivoting, and explicitly request order where needed.
Spilling to disk could also help workloads that exceed available memory, but its benefits are limited to the operations currently supported. The stated 80% RAM threshold and 64 GB disk budget are defaults, not guarantees that every query will complete under memory pressure. SQL support and the new Map dtype may make Polars usable for a wider range of data workflows, though the actual fit depends on query patterns and compatibility requirements.
Polars reports that it led competing engines on most of its tested SQL benchmarks. That is relevant to users evaluating query speed, but the figures come from tests designed and run by the Polars team. They should be treated as vendor-reported comparisons rather than a universal ranking of database engines.
high performance data processing laptop
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Polars Tested SQL Performance
Polars compared its SQL engine with DuckDB 1.5.6, DuckDB 2.0 alpha and DataFusion 54.0.0 using data derived from TPC-H and TPC-DS. The team ran tests on an AWS c7a.4xlarge instance with 16 virtual CPUs and 32 GB of memory, and a c7a.metal instance with 192 virtual CPUs and 384 GB.
According to the report, each query ran five times in a hot setting, with the best run used for comparisons. The file cache was cleared between engines and benchmark suites, but not between queries. Polars said its default configuration was fastest on all but one benchmark, while also reporting overhead on the 192-thread machine that hurt smaller queries. It said a 32-core Polars configuration was competitive or ahead across the benchmarks. The project published code for replication, and says the large-machine overhead has been diagnosed, with a fix hoped for in a future release.
The report also notes limits in the results: DataFusion timed out on TPC-DS query 72, timed out once on query 67, and ran out of memory on TPC-H query 18 on the smaller machine. Those queries were excluded from the results for every engine. The release post says Polars and both DuckDB versions completed all queries, but benchmark exclusions and the specific hardware and settings matter when interpreting the comparison.
“Calling collect on a LazyFrame will now default to the streaming engine.”
— Polars, in its version 2.0 release report
As an affiliate, we earn on qualifying purchases.
Limits of the New Defaults
The release report does not specify when it was published, and the supplied material does not give a separate independent review of the release or its benchmarks. The performance results therefore remain Polars-reported measurements; results for other workloads, machines and configurations may differ.
It is also unclear from the report how the approximate 80% memory threshold will behave across different systems, or how much disk use particular workloads may require. Out-of-core support for joins and group-by operations is described as future work, without a delivery date. Users should also verify how their own queries behave under the new streaming default, especially where row order is significant.
professional data analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
More Spill Support Planned
Polars says it plans to extend out-of-core processing to joins and group-by operations, but has not given a schedule in the supplied release report. The team also says it has diagnosed the overhead seen when running at 192 threads and hopes to address it in a later release.
For now, users can test version 2.0 against their own workloads, check ordering assumptions and review disk and memory behavior for supported operations. Polars has shared a benchmark repository for readers who want to reproduce its SQL comparisons. Independent replication and additional release details would help establish how broadly the reported performance results apply.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main change in Polars 2.0?
Collecting a LazyFrame now uses the streaming engine by default. Polars says this can improve memory use and performance on many queries, while warning that some operations do not preserve row order automatically.
Does Polars 2.0 spill every query to disk?
No. The release enables spill-to-disk support for selected operations, including sorts, window functions and many expressions. Polars says support for joins and group-by operations is planned, not yet included.
How should users handle row order after upgrading?
Check whether an application relies on row order for joins, group-by operations or unpivoting. Polars says users can request preserved order for supported operations with maintain_order=True.
Are the SQL benchmark results independent?
No independent benchmark is described in the supplied material. The comparisons were conducted and reported by Polars, using specified TPC-H and TPC-DS-derived workloads, hardware and test settings.
What is the new Map dtype for?
The Map dtype represents Arrow MapType data directly in Polars. The release report describes dictionary-like operations, including retrieving values by key, checking whether a key exists, and accessing keys or values.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
