Skip to content

Version 2.0-rc

Breaking changes

The Lazy API defaults to the streaming engine

LazyFrame.collect and collect_async still default to engine="auto", but "auto" now resolves to the streaming engine for lazy queries (it used to resolve to the in-memory engine). Eager DataFrame operations are unaffected: they continue to resolve "auto" to the in-memory engine internally.

Note that sink_* is not affected by this change: writing to a file was already dispatched to the streaming engine regardless of engine.

explain() and show_graph() are not affected either: they only render a streaming plan when engine="streaming" is passed explicitly (see below).

Danger

This change may silently impact the results of your pipelines. The streaming engine does not guarantee row order for operations that don't require it (unpivot, group_by, joins, ...). If your code relies on incidental ordering, sort explicitly, or pass maintain_order where the operation supports it (e.g. join(..., maintain_order="left")).

Example

Before:

>>> lf = pl.LazyFrame({"a": ["x", "y", "z"], "b": [1, 2, 3], "c": [4, 5, 6]})
>>> lf.unpivot(pl.selectors.numeric(), index="a").collect()
shape: (6, 3)
┌─────┬──────────┬───────┐
│ a   ┆ variable ┆ value │
│ --- ┆ ---      ┆ ---   │
│ str ┆ str      ┆ i64   │
╞═════╪══════════╪═══════╡
│ x   ┆ b        ┆ 1     │
│ y   ┆ b        ┆ 2     │
│ z   ┆ b        ┆ 3     │
│ x   ┆ c        ┆ 4     │
│ y   ┆ c        ┆ 5     │
│ z   ┆ c        ┆ 6     │
└─────┴──────────┴───────┘

After:

>>> lf.unpivot(pl.selectors.numeric(), index="a").collect()
shape: (6, 3)
┌─────┬──────────┬───────┐
│ a   ┆ variable ┆ value │
│ --- ┆ ---      ┆ ---   │
│ str ┆ str      ┆ i64   │
╞═════╪══════════╪═══════╡
│ x   ┆ b        ┆ 1     │
│ x   ┆ c        ┆ 4     │
│ y   ┆ b        ┆ 2     │
│ y   ┆ c        ┆ 5     │
│ z   ┆ b        ┆ 3     │
│ z   ┆ c        ┆ 6     │
└─────┴──────────┴───────┘

The exact row order shown above is not guaranteed — only that it is no longer the pre-2.0 order. Use instead, if you rely on the order:

>>> lf.unpivot(pl.selectors.numeric(), index="a").collect().sort(pl.all())

Joins are affected the same way, and are easy to miss since nothing about the query looks order-sensitive:

>>> left = pl.LazyFrame({"k": [0, 1, 2], "l": ["a", "b", "c"]})
>>> right = pl.LazyFrame({"k": [2, 1, 0], "r": ["x", "y", "z"]})
>>> left.join(right, on="k", how="left").collect()  # row order no longer matches `left`

Use instead, if you rely on the left-hand row order:

>>> left.join(right, on="k", how="left", maintain_order="left").collect()

To restore the previous default engine altogether:

>>> pl.Config.set_engine_affinity("in-memory")  # process-wide
>>> lf.collect(engine="in-memory")  # per query

Or set the POLARS_ENGINE_AFFINITY=in-memory environment variable.

SQL is affected too. SQLContext.execute(eager=True) and pl.sql(..., eager=True) used to collect through the in-memory engine; they now go through LazyFrame.collect() like any other lazy query.

>>> pl.sql("SELECT * FROM t WHERE k > 1", eager=True)  # now runs on the streaming engine

To keep using the in-memory engine for a SQL query, leave the result lazy with eager=False and pick the engine when you collect it yourself:

>>> pl.sql("SELECT * FROM t WHERE k > 1", eager=False).collect(engine="in-memory")

The engine affinity settings above apply to SQL as well.

pl.read_csv is now dispatched to pl.scan_csv(...).collect().

It gains scan_csv's with_column_names, infer_schema_files, credential_provider, include_file_paths, extra_columns, and missing_columns parameters, and accepts a list of sources.

It loses n_threads, batch_size, sample_size, and rechunk, which have no equivalent in the lazy reader (call .rechunk() on the result if you need it). Its behavior now also matches scan_csv exactly: a schema_overrides list must cover every column in the file, and columns=[...] returns columns in the requested order rather than sorted.

Example

Before:

>>> data = b"a,b,c,d\n1,2,3,4\n1,2,3,4\n"
>>> pl.read_csv(data, schema_overrides=[pl.String]).dtypes
[String, Int64, Int64, Int64]
>>> pl.read_csv(data, columns=[2, 1, 3]).columns
['b', 'c', 'd']

After:

>>> pl.read_csv(data, schema_overrides=[pl.String])
Traceback (most recent call last):
...
polars.exceptions.SchemaError: The number of dtypes in schema override must be equal to the number of fields in the file (1 != 4).
>>> pl.read_csv(data, columns=[2, 1, 3]).columns
['c', 'b', 'd']

Use instead:

>>> pl.read_csv(data, schema_overrides=[pl.String, pl.Int64, pl.Int64, pl.Int64]).dtypes
[String, Int64, Int64, Int64]

pl.read_ipc is now dispatched to pl.scan_ipc(...).collect() for all non-use_pyarrow inputs

The memory_map and rechunk parameters are removed. Unlike scan_ipc's memory_map, which had already become a no-op in 1.40, read_ipc's memory_map had real effect before this release.

Example

Before:

>>> pl.read_ipc("data.arrow", memory_map=True)

After:

>>> pl.read_ipc("data.arrow", memory_map=True)
Traceback (most recent call last):
...
TypeError: read_ipc() got an unexpected keyword argument 'memory_map'

scan_ipc loses the same memory_map parameter in this release (it had been a no-op there since 1.40; see the removal table below), and both readers also lose rechunk. Use .rechunk() on the returned frame instead.

Use instead:

>>> pl.read_ipc("data.arrow")

Change default show_graph() plan_stage to "physical"

LazyFrame.show_graph()'s plan_stage parameter now defaults to "physical" instead of "ir". This does not, by itself, change what the default call renders: plan_stage="physical" only draws the streaming physical graph when engine="streaming" is also selected (explicitly, or via engine affinity); with the default engine="auto" it still falls back to the optimized IR graph, same as before.

Example

Before:

>>> lf.group_by("a").agg(pl.all().sum()).show_graph()  # shows the optimized IR plan

After:

>>> lf.group_by("a").agg(pl.all().sum()).show_graph()  # plan_stage="physical" by default, but still
>>> #                                                     shows the IR plan, since engine="auto" != "streaming"
>>> lf.group_by("a").agg(pl.all().sum()).show_graph(engine="streaming")  # shows the streaming physical plan

Use instead, to restore the previous default:

>>> lf.group_by("a").agg(pl.all().sum()).show_graph(plan_stage="ir")

Remove LazyFrame.profile()

LazyFrame.profile() has been removed. It was designed for the in-memory engine; the concurrent nature of the streaming engine, now the default, would make its per-node timings misleading.

Example

Before:

>>> lf.group_by("a", maintain_order=True).agg(pl.all().sum()).sort("a").profile()
(shape: (3, 3)
┌─────┬─────┬─────┐
│ a   ┆ b   ┆ c   │
│ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ i64 │
╞═════╪═════╪═════╡
│ x   ┆ 1   ┆ 4   │
│ y   ┆ 2   ┆ 5   │
│ z   ┆ 3   ┆ 6   │
└─────┴─────┴─────┘, shape: (2, 3)
┌──────────────┬───────┬──────┐
│ node         ┆ start ┆ end  │
│ ---          ┆ ---   ┆ ---  │
│ str          ┆ u64   ┆ u64  │
╞══════════════╪═══════╪══════╡
│ optimization ┆ 0     ┆ 371  │
│ sort(a)      ┆ 1162  ┆ 1234 │
└──────────────┴───────┴──────┘)

After:

>>> lf.group_by("a", maintain_order=True).agg(pl.all().sum()).sort("a").profile()
Traceback (most recent call last):
...
polars.exceptions.AttributeRemovedError: `profile` was removed in version 2.0; It was designed for the in-memory engine and would give misleading per-node timings under the streaming engine (now the default).

Users of Polars Cloud / On-Prem can use the Query Profiler functionality for detailed metrics. We're working on a way to provide query profiling for open source Polars.

Update the strict behavior of pl.concat()/pl.union()

how="horizontal" now always requires equal heights, matching what strict=True used to mean. Previously, the default silently padded shorter frames with null.

Danger

This change may silently impact the results of your pipelines. pl.concat([df1, df2], how="horizontal") used to pad; now it raises if the heights differ.

Example

Before:

>>> df1 = pl.DataFrame({"a": [1, 2, 3]})
>>> df2 = pl.DataFrame({"b": [4, 5]})
>>> pl.concat([df1, df2], how="horizontal")
shape: (3, 2)
┌─────┬──────┐
│ a   ┆ b    │
│ --- ┆ ---  │
│ i64 ┆ i64  │
╞═════╪══════╡
│ 1   ┆ 4    │
│ 2   ┆ 5    │
│ 3   ┆ null │
└─────┴──────┘

After:

>>> pl.concat([df1, df2], how="horizontal")
Traceback (most recent call last):
...
polars.exceptions.ShapeError: cannot concat dataframes with different heights in 'strict' mode

Use instead, padding the shorter frame yourself so that the intent is explicit at the call site:

>>> df2_padded = df2.select(pl.col("b").extend_constant(None, df1.height - df2.height))
>>> pl.concat([df1, df2_padded], how="horizontal")
shape: (3, 2)
┌─────┬──────┐
│ a   ┆ b    │
│ --- ┆ ---  │
│ i64 ┆ i64  │
╞═════╪══════╡
│ 1   ┆ 4    │
│ 2   ┆ 5    │
│ 3   ┆ null │
└─────┴──────┘

If you do want Polars to pad for you, how="horizontal_extend" is the direct equivalent of the old behavior:

>>> pl.concat([df1, df2], how="horizontal_extend")  # same result as above

strict=False is no longer accepted as a way to opt into padding either. It now raises, since how="horizontal" itself no longer pads.

>>> pl.concat([df1, df2], how="horizontal", strict=False)
Traceback (most recent call last):
...
ValueError: `strict=False` is no longer supported for `how='horizontal'`. Use `how='horizontal_extend'` to pad shorter frames with `null`.

strict cannot be combined with how="horizontal_extend" at all:

>>> pl.concat([df1, df2], how="horizontal_extend", strict=True)
Traceback (most recent call last):
...
ValueError: `strict` cannot be used with `how='horizontal_extend'`

Set explode() empty_as_null=False by default

An empty list now explodes into zero rows, rather than one null row. keep_nulls is unaffected: a null list still explodes into one null row.

Danger

This change may silently impact the results of your pipelines. Row counts change wherever a list column contains empty lists.

Example

Before:

>>> df = pl.DataFrame({"a": [[1, 2, 3], [], [4, 5, 6]]})
>>> df.explode("a")
shape: (7, 1)
┌──────┐
│ a    │
│ ---  │
│ i64  │
╞══════╡
│ 1    │
│ 2    │
│ 3    │
│ null │
│ 4    │
│ 5    │
│ 6    │
└──────┘

After:

>>> df.explode("a")
shape: (6, 1)
┌─────┐
│ a   │
│ --- │
│ i64 │
╞═════╡
│ 1   │
│ 2   │
│ 3   │
│ 4   │
│ 5   │
│ 6   │
└─────┘

Use instead:

>>> df.explode("a", empty_as_null=True)

Change supertype of signed integer types and UInt64 from Float64 to Int128

Adding (or otherwise combining) a signed integer column with a UInt64 column used to produce a lossy Float64 result. It now produces an exact Int128 result.

Danger

This change may silently impact the results of your pipelines. No error is raised: both the output dtype and the computed values change.

Example

Before:

>>> lf = pl.LazyFrame({"a": [1, 2, 3], "b": [1, 2, 3]}, schema={"a": pl.Int64, "b": pl.UInt64})
>>> lf.select((pl.col("a") + pl.col("b")).alias("result")).collect_schema()
Schema({'result': Float64})

After:

>>> lf.select((pl.col("a") + pl.col("b")).alias("result")).collect_schema()
Schema({'result': Int128})

Make coercion casts for is_in() strict instead of lossy

is_in() no longer lossily coerces operands to a shared supertype; it now only coerces when doing so is lossless, and raises otherwise.

Example

Before:

>>> pl.Series([1]).is_in(pl.Series([1.99]))
shape: (1,)
Series: '' [bool]
[
        false
]

After:

>>> pl.Series([1]).is_in(pl.Series([1.99]))
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: 'is_in' cannot check for Int64 values in List(Float64) data.
Hint: Before version 2.0, Polars would perform this check by lossily coercing the operands to Float64. However, since Polars 2.0, for is_in() it is required to explicitly cast (one of) the operands to a compatible type.

Use instead:

>>> pl.Series([1]).is_in(pl.Series([1.99]).cast(pl.Int64))

Stop coercing pl.col(...) to a selector in selector &, |, ^

Combining a selector with pl.col(...) via &, |, or ^ used to be special-cased to mean a set operation on the selected columns. It now falls through to the ordinary element-wise operator.

Danger

This change may silently impact the results of your pipelines. It only raises when the involved dtypes are incompatible; for compatible dtypes (e.g. two integer columns) it silently switches from a column selection to an element-wise bitwise operation.

Example

Before:

>>> df = pl.DataFrame({"a": [1, 2], "b": [3, 4], "mask": [1, 0]})
>>> df.select(pl.selectors.integer() & pl.col("mask"))
shape: (2, 1)
┌──────┐
│ mask │
│ ---  │
│ i64  │
╞══════╡
│ 1    │
│ 0    │
└──────┘

After:

>>> df.select(pl.selectors.integer() & pl.col("mask"))
shape: (2, 3)
┌─────┬─────┬──────┐
│ a   ┆ b   ┆ mask │
│ --- ┆ --- ┆ ---  │
│ i64 ┆ i64 ┆ i64  │
╞═════╪═════╪══════╡
│ 1   ┆ 1   ┆ 1    │
│ 0   ┆ 0   ┆ 0    │
└─────┴─────┴──────┘

Note that no error is raised: a, b, and mask (every column selected by integer()) are each combined with mask via a bitwise &, element-wise, instead of the query selecting just the mask column.

Use instead:

>>> df.select(pl.selectors.integer() & pl.selectors.by_name("mask"))

Set the output name of pl.datetime and pl.repeat to their leftmost argument name

pl.datetime(...) and pl.repeat(...) used to always name their output column "datetime" / "repeat". For consistency with other multi-argument functions, they now take the name of their leftmost argument (or "literal", if that argument is itself a literal).

Danger

This change may silently impact the results of your pipelines. No error is raised. Only the output column name changes, which can silently overwrite an existing column in a with_columns call, or break a df["datetime"] lookup.

Example

Before:

>>> df = pl.DataFrame({"year": [2001], "month": [1], "day": [1], "hour": [23]})
>>> df.select(pl.datetime("year", "month", "day", "hour")).columns
['datetime']

After:

>>> df.select(pl.datetime("year", "month", "day", "hour")).columns
['year']

Use instead, if you rely on the name:

>>> df.select(pl.datetime("year", "month", "day", "hour").alias("datetime")).columns
['datetime']

Similarly for pl.repeat:

>>> pl.repeat(3, n=3, dtype=pl.Int8, eager=True)
# before: Series: 'repeat' [i8]
# after:  Series: 'literal' [i8]

Preserve outer nulls in list/array to_struct()

A null list or array row now converts to a null struct, rather than a struct whose fields are all null.

Danger

This change may silently impact the results of your pipelines. No error is raised: row validity changes, which affects is_null(), joins, and to_list().

Example

Before:

>>> s = pl.Series([None], dtype=pl.List(pl.Int64))
>>> s.list.to_struct(fields=["a", "b", "c"]).to_list()
[{'a': None, 'b': None, 'c': None}]

After:

>>> s.list.to_struct(fields=["a", "b", "c"]).to_list()
[None]

Do not reset seek position when scanning from file-like objects

Reading from a file-like object (e.g. io.BytesIO) no longer implicitly rewinds it to the start. Scanning now starts from the object's current position.

Danger

The common buf = io.BytesIO(); df.write_parquet(buf); pl.read_parquet(buf) round trip now needs an explicit buf.seek(0) between the write and the read.

Example

Before:

>>> import io
>>> buf = io.BytesIO()
>>> pl.DataFrame({"a": [1, 2, 3]}).write_parquet(buf)
>>> pl.read_parquet(buf)
shape: (3, 1)
┌─────┐
│ a   │
│ --- │
│ i64 │
╞═════╡
│ 1   │
│ 2   │
│ 3   │
└─────┘

After:

>>> pl.read_parquet(buf)
Traceback (most recent call last):
...
polars.exceptions.ComputeError: parquet: File out of specification: A Parquet file must contain a header and footer with at least 12 bytes

Use instead:

>>> buf.seek(0)
>>> pl.read_parquet(buf)

Preserve height in zero-width DataFrame/LazyFrame operations

DataFrame/LazyFrame with zero columns now carry a height rather than collapsing to (0, 0). This affects drop() of every column, gather_every() on zero-width frames, and manually constructed empty frames: pl.DataFrame() and pl.LazyFrame() now have a fixed height of 0, so adding a longer column via with_columns() raises instead of adopting the new column's length.

Example

Before:

>>> df = pl.DataFrame({"a": [2, 1, 3], "b": ["a", "b", "c"], "c": [1, 2, 3]})
>>> df.drop("*").shape
(0, 0)
>>> pl.DataFrame().with_columns(pl.Series([None, None])).shape
(2, 1)

After:

>>> df.drop("*").shape
(3, 0)
>>> pl.DataFrame().with_columns(pl.Series([None, None]))
Traceback (most recent call last):
...
polars.exceptions.ShapeError: can't broadcast Series '' of length 2 to length 0

Use instead, to build a frame from scratch:

>>> pl.DataFrame(height=2).with_columns(pl.Series([None, None])).shape
(2, 1)

Always return Series from from_arrow(<ArrowStreamExportable>)

pl.from_arrow() on an object that supports only the Arrow PyCapsule stream interface (i.e. __arrow_c_stream__, but not a pyarrow.Table/RecordBatch itself) now always returns a Series, rather than a DataFrame. A multi-column source becomes a Series of Struct, one struct per row. The schema/schema_overrides parameters, which had no effect on this path, now raise instead of being silently ignored.

Danger

This change may silently impact the results of your pipelines: the return type changes from DataFrame to Series, and no error is raised.

Example

Before:

>>> class ArrowStreamOnly:
...     """Exposes only `__arrow_c_stream__`, like a third-party Arrow producer."""
...     def __init__(self, obj):
...         self._obj = obj
...     def __arrow_c_stream__(self, requested_schema=None):
...         return self._obj.__arrow_c_stream__(requested_schema)
>>> pl.from_arrow(ArrowStreamOnly(pl.DataFrame({"x": [1, 2, 3]})))
shape: (3, 1)
┌─────┐
│ x   │
│ --- │
│ i64 │
╞═════╡
│ 1   │
│ 2   │
│ 3   │
└─────┘

After:

>>> pl.from_arrow(ArrowStreamOnly(pl.DataFrame({"x": [1, 2, 3]})))
shape: (3,)
Series: '' [struct[1]]
[
        {1}
        {2}
        {3}
]

Use instead, if you need a DataFrame:

>>> pl.from_arrow(ArrowStreamOnly(pl.DataFrame({"x": [1, 2, 3]}))).struct.unnest()

Remove the DataFrame Interchange Protocol

Polars no longer implements the DataFrame Interchange Protocol. DataFrame.__dataframe__() has been removed, along with the polars.interchange.buffer, column, dataframe, from_dataframe, and utils submodules and the polars.interchange.protocol symbols (DtypeKind, Buffer, Column, SupportsInterchange, CopyNotAllowedError, ...). The polars.interchange package itself remains, but now holds only CompatLevel.

pl.from_dataframe() is kept, but only accepts objects supporting the Arrow PyCapsule interface; it no longer falls back to the interchange protocol. Its allow_copy parameter has been removed.

Example

Before:

>>> df = pl.DataFrame({"a": [1, 2, 3]})
>>> df.__dataframe__()
<polars.interchange.dataframe.PolarsDataFrame object at 0x10c0f5d90>

After:

>>> df.__dataframe__()
Traceback (most recent call last):
...
polars.exceptions.AttributeRemovedError: `__dataframe__` was removed in version 2.0; the dataframe interchange protocol is not supported anymore. Consider using `to_arrow` or `to_pandas` instead.

Use instead, for a consumer that accepted an interchange object (e.g. Seaborn):

>>> import seaborn as sns
>>> sns.scatterplot(df.to_pandas(), x="a", y="a")

For Arrow-native interop, use DataFrame.to_arrow() and pl.from_arrow().

Passing an object that supports neither interface to pl.from_dataframe() now raises, rather than silently going through the interchange protocol:

>>> pl.from_dataframe(object())
Traceback (most recent call last):
...
TypeError: expected object supporting the PyCapsule Interface, got 'object'

And allow_copy is gone:

>>> pl.from_dataframe(df, allow_copy=True)
Traceback (most recent call last):
...
polars.exceptions.ArgumentRemovedError: the argument 'allow_copy' for 'from_dataframe' was deprecated in version 1.23.0 and has been removed in version 2.0.0.

Allow transposing an empty DataFrame

DataFrame.transpose() on a zero-height or zero-width DataFrame no longer raises; it transposes the frame's shape like any other, consistent with transpose() on a non-empty frame.

Example

Before:

>>> pl.DataFrame().transpose()
Traceback (most recent call last):
...
polars.exceptions.NoDataError: unable to transpose an empty DataFrame

After:

>>> pl.DataFrame().transpose()
shape: (0, 0)
┌┐
╞╡
└┘

Note that a frame's height becomes the transposed frame's width, and vice versa, so a zero-width frame with a nonzero number of rows transposes into a frame with that many columns and no rows. The output columns get the usual generated names, and Null dtype, since there is no input data to derive a supertype from:

>>> pl.DataFrame(height=3).transpose()
shape: (0, 3)
┌──────────┬──────────┬──────────┐
│ column_0 ┆ column_1 ┆ column_2 │
│ ---      ┆ ---      ┆ ---      │
│ null     ┆ null     ┆ null     │
╞══════════╪══════════╪══════════╡
└──────────┴──────────┴──────────┘

The reverse case, a zero-height frame with columns, drops those columns and keeps their count as the new height:

>>> pl.DataFrame(schema={"a": pl.Int32, "b": pl.Int32}).transpose().shape
(2, 0)
>>> pl.DataFrame(schema={"a": pl.Int32, "b": pl.Int32}).transpose(include_header=True)
shape: (2, 1)
┌────────┐
│ column │
│ ---    │
│ str    │
╞════════╡
│ a      │
│ b      │
└────────┘

shift(n=None) now raises

A null shift amount used to emit a DeprecationWarning and return a column of nulls. It is now a hard error. This applies to DataFrame.shift, LazyFrame.shift, and Expr.shift, with or without fill_value.

Example

Before:

>>> df = pl.DataFrame({"a": pl.Series([1, 2, 3], dtype=pl.Int32)})
>>> df.shift(None)  # DeprecationWarning
shape: (3, 1)
┌──────┐
│ a    │
│ ---  │
│ i32  │
╞══════╡
│ null │
│ null │
│ null │
└──────┘

After:

>>> df.shift(None)
Traceback (most recent call last):
...
polars.exceptions.ComputeError: shift value 'n' must not be null.

This error occurred in the following expression:
    col("a").shift([null])

Use instead:

>>> df.shift(1)
shape: (3, 1)
┌──────┐
│ a    │
│ ---  │
│ i32  │
╞══════╡
│ null │
│ 1    │
│ 2    │
└──────┘

Unknown Arrow extension types load as extension dtypes

When Polars reads Arrow data carrying an ARROW:extension:name it does not recognize, it used to warn once and load the column as its storage type. It now loads it as a generic pl.Extension(name, storage) dtype, and the warning is gone. The pre-2.0 warning announced this change ("In Polars 2.0, the default behavior will change to 'load_as_extension'").

Danger

This change may silently impact the results of your pipelines. No error is raised: the schema of anything read from Arrow with extension metadata changes, which affects dtype comparisons, casts, and any operation not defined on Extension.

Example

Before:

>>> import pyarrow as pa
>>> field = pa.field(
...     "id", pa.int64(), metadata={b"ARROW:extension:name": b"google:sqlType:integer"}
... )
>>> tbl = pa.table([pa.array([1, 2, 3])], schema=pa.schema([field]))
>>> pl.from_arrow(tbl).schema  # warns: Extension type '...' is not registered
Schema({'id': Int64})

After:

>>> pl.from_arrow(tbl).schema
Schema({'id': Extension('google:sqlType:integer', Int64)})

Use instead, to recover the storage dtype:

>>> pl.from_arrow(tbl).select(pl.all().ext.storage()).schema
Schema({'id': Int64})

Or restore the previous behavior process-wide by setting the POLARS_UNKNOWN_EXTENSION_TYPE_BEHAVIOR=load_as_storage environment variable.

Casts and operations no longer supported

The following casts and operations, previously deprecated, now raise an error.

Disable casting from integers to categoricals, and from categoricals to integers

Use .cat.to() / .cat.physical() instead.

Example

Before:

>>> dtype = pl.Enum(["a", "b", "c"])
>>> pl.Series([None, 1, 0, 2], dtype=pl.UInt32).cast(dtype)
shape: (4,)
Series: '' [enum]
[
        null
        "b"
        "a"
        "c"
]

After:

>>> pl.Series([None, 1, 0, 2], dtype=pl.UInt32).cast(dtype)
Traceback (most recent call last):
...
polars.exceptions.ComputeError: casting from u32 to enum is not supported.
Instead of `.cast(Enum([...])`, use `.cat.to(Enum([...]))`.

Use instead:

>>> pl.Series([None, 1, 0, 2], dtype=pl.UInt32).cat.to(dtype)

And the reverse:

>>> s = pl.Series("a", ["cat2", "cat0", "cat1"], dtype=pl.Enum(["cat0", "cat1", "cat2"]))
>>> s.cast(pl.UInt32)
Traceback (most recent call last):
...
polars.exceptions.ComputeError: cannot cast categorical types to UInt32.
Instead of `.cast(UInt32)`, use `.cat.physical()`.
>>> s.cat.physical()  # instead

Remove casts from string to temporal types

Casting a String column directly to Date/Datetime/Time is no longer supported.

Example

Before:

>>> pl.Series(["2022-08-30"]).cast(pl.Date)
shape: (1,)
Series: '' [date]
[
        2022-08-30
]

After:

>>> pl.Series(["2022-08-30"]).cast(pl.Date)
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: casting from string to date is not supported.
It was removed in Polars 2.0. Use `str.to_date()` instead.

Use instead:

>>> pl.Series(["2022-08-30"]).str.to_date()

The same applies to casting String to Datetime and Time — use str.to_datetime() and str.to_time() instead.

Disallow casting from non-nested types to pl.List(..)

Example

Before:

>>> pl.Series("a", [1, 2, 3], dtype=pl.Int32).cast(pl.List(pl.Int32))
shape: (3,)
Series: 'a' [list[i32]]
[
        [1]
        [2]
        [3]
]

After:

>>> pl.Series("a", [1, 2, 3], dtype=pl.Int32).cast(pl.List(pl.Int32))
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: casting from Int32 to list type is not supported
Hint: Use pl.list(expr) to turn the Int32 column into a column of single-element lists.

Use instead:

>>> pl.select(pl.list(pl.Series("a", [1, 2, 3], dtype=pl.Int32)))

Do not allow boolean operators between booleans and integer types

Example

Before:

>>> lf = pl.LazyFrame({"bool": [True, False], "int": [1, 2]}, schema={"bool": pl.Boolean, "int": pl.Int32})
>>> lf.select(pl.col("bool") & pl.col("int")).collect_schema()
Schema({'bool': Int32})

After:

>>> lf.select(pl.col("bool") & pl.col("int")).collect_schema()
Traceback (most recent call last):
...
polars.exceptions.ComputeError: & on Boolean and Int32 is not supported
Hint: cast the Boolean to Int32 using pl.Expr.cast().

Use instead:

>>> lf.select(pl.col("bool").cast(pl.Int32) & pl.col("int")).collect_schema()

Remove support for std()/var() on the Duration dtype

Example

Before:

>>> from datetime import timedelta
>>> pl.Series([timedelta(days=1), timedelta(days=2), timedelta(days=4)]).std()
datetime.timedelta(days=1, seconds=45578, microseconds=180014)

After:

>>> pl.Series([timedelta(days=1), timedelta(days=2), timedelta(days=4)]).std()
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: `std` operation not supported for dtype `duration[μs]`

This also applies to var(), ewm_std(), and ewm_var() on Duration columns. Cast to an integer representation first if you need the statistic, e.g. .dt.total_microseconds().std().

Disable invalid struct casts when strict=True

Casting between two Struct dtypes with a different number of fields, or with mismatched field names, used to silently succeed by taking the first n matching fields. The same kind of silent truncation struct.rename_fields() used to do (see below). Under the default strict=True it now raises; pass strict=False to keep the old truncating behavior.

Danger

This change may silently impact the results of your pipelines: casts that used to silently drop struct fields now raise, unless you pass strict=False.

Example

Before:

>>> s = pl.Series("x", [{"a": 1, "b": 2}])
>>> s.cast(pl.Struct({"a": pl.Int64}))
shape: (1,)
Series: 'x' [struct[1]]
[
        {1}
]

After:

>>> s.cast(pl.Struct({"a": pl.Int64}))
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: cast from `struct[2]` to `struct[1]` failed in column 'x': structs do not have the same number of fields: 2 vs 1

Ensure that any output struct has the same number of fields as the input, and that all struct field names in the output are present in the input.
Use `strict=False` to force the cast, and Polars will select the first n fields from the struct.

Use instead, to keep the previous behavior:

>>> s.cast(pl.Struct({"a": pl.Int64}), strict=False)

Disallow flat dtypes where a list is expected

list.gather, str.contains_any, and str.replace_many used to accept a flat (non-list) argument where a list is expected, implicitly imploding it. That coercion has warned since it was deprecated (see #22149) and existed only as a pre-2.0 compatibility shim; it has now been removed.

Example

Before, each of these silently imploded its flat argument:

>>> df = pl.DataFrame({"a": [[1, 2], [3, 4]], "b": [0, 1]})
>>> df.select(pl.col("a").list.gather(0))                  # 0 imploded to [0]
>>> df.select(pl.col("a").list.gather(pl.col("b")))        # all of `b` imploded to one list
>>> pl.select(pl.lit("ab").str.contains_any(pl.lit("a")))  # "a" imploded to ["a"]

After:

>>> df.select(pl.col("a").list.gather(0))
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: `list.gather` indices must be a list of integers, not a flat dyn int. Use `implode` to wrap the flat value into a list.
>>> pl.select(pl.lit("ab").str.contains_any(pl.lit("a")))
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: `str.contains_any` with a flat string datatype is invalid. Use `implode` to wrap the string value into a list.

Use an explicit implode instead; it reproduces the old behavior exactly, which is what the error message suggests:

>>> df.select(pl.col("a").list.gather([0]))                    # -> [1], [3]
>>> df.select(pl.col("a").list.gather(pl.col("b").implode()))  # -> [1, 2], [3, 4]
>>> pl.select(pl.lit("ab").str.contains_any(["a"]))

Note that the flat-column form never indexed per row: list.gather(pl.col("b")) imploded all of b into a single index list, so every row was gathered with the same indices. If per-row indexing is what you want, use list.get, or concat_list to keep a list-typed result:

>>> df.select(pl.col("a").list.get(pl.col("b")))                     # -> 1, 4
>>> df.select(pl.col("a").list.gather(pl.concat_list(pl.col("b"))))  # -> [1], [4]

API reshapes

Simplify list.to_struct API

fields is now a required, positional argument. The n_field_strategy, upper_bound parameters and the callable form of fields have been removed; the number of names given determines the output width.

Example

Before:

>>> df = pl.DataFrame({"n": [[0, 1, 2], [0, 1]]})
>>> df.select(pl.col("n").list.to_struct(upper_bound=3))
shape: (2, 1)
┌────────────┐
│ n          │
│ ---        │
│ struct[3]  │
╞════════════╡
│ {0,1,2}    │
│ {0,1,null} │
└────────────┘

After:

>>> df.select(pl.col("n").list.to_struct())
Traceback (most recent call last):
...
TypeError: ExprListNameSpace.to_struct() missing 1 required positional argument: 'fields'

Use instead:

>>> df.select(pl.col("n").list.to_struct(["field_0", "field_1", "field_2"]))

Both removed parameters raise a typed error with a pointer to fields:

>>> df.select(pl.col("n").list.to_struct(upper_bound=3))
Traceback (most recent call last):
...
polars.exceptions.ArgumentRemovedError: the argument 'upper_bound' for 'ExprListNameSpace.to_struct' has been removed in version 2.0. Pass the field names explicitly via `fields` instead, e.g. `fields=[f"field_{i}" for i in range(upper_bound)]`.

Note that arr.to_struct() (for the fixed-size Array dtype) still accepts an omitted fields, since the output width is already known from the dtype; only the callable form of fields was removed there.

Make add_business_days() and str.to_decimal() keyword-only arguments strict

Expr/Series.dt.add_business_days() (all arguments after n) and Expr/Series.str.to_decimal() (inference_length) have accepted these as keyword-only for several releases, with a DeprecationWarning if passed positionally. That warning is now a hard error.

Example

Before:

>>> import datetime as dt
>>> week_mask = (True, True, True, True, True, False, False)
>>> pl.Series([dt.date(2020, 1, 1)]).dt.add_business_days(5, week_mask)
shape: (1,)
Series: '' [date]
[
        2020-01-08
]
>>> pl.Series(["1.5"]).str.to_decimal(100)
shape: (1,)
Series: '' [decimal[2,1]]
[
        1.5
]

After:

>>> pl.Series([dt.date(2020, 1, 1)]).dt.add_business_days(5, week_mask)
Traceback (most recent call last):
...
TypeError: ExprDateTimeNameSpace.add_business_days() takes 2 positional arguments but 3 were given
>>> pl.Series(["1.5"]).str.to_decimal(100)
Traceback (most recent call last):
...
TypeError: StringNameSpace.to_decimal() takes 1 positional argument but 2 were given

Use instead:

>>> pl.Series([dt.date(2020, 1, 1)]).dt.add_business_days(5, week_mask=week_mask)
>>> pl.Series(["1.5"]).str.to_decimal(inference_length=100)

Replace multi-seed hash API with a single seed

Expr.hash, Series.hash, and DataFrame.hash_rows no longer accept seed_1/seed_2/seed_3.

Hash values computed for the default seed also changed. Note that Polars doesn't guarantee hash stability across versions, only within a single version.

Example

Before:

>>> df = pl.DataFrame({"a": [1, 2, 3]})
>>> df.with_columns(pl.all().hash(10, 20, 30, 40))

After:

>>> df.with_columns(pl.all().hash(10, 20, 30, 40))
Traceback (most recent call last):
...
TypeError: Expr.hash() takes from 1 to 2 positional arguments but 5 were given

Use instead:

>>> df.with_columns(pl.all().hash(10))

Do not support list type anymore for Series.search_sorted()

Passing a Python list was ambiguous. It could mean either a single list-typed search value, or multiple scalar search targets. It now raises.

Example

Before:

>>> pl.Series([1, 2, 3]).search_sorted([1, 2])
shape: (2,)
Series: '' [u32]
[
        0
        1
]

After:

>>> pl.Series([1, 2, 3]).search_sorted([1, 2])
Traceback (most recent call last):
...
polars.exceptions.InvalidOperationError: passing a list to `search_sorted` is ambiguous; use `pl.Series([...])` or `pl.lit(...)`

Use instead:

>>> pl.Series([1, 2, 3]).search_sorted(pl.Series([1, 2]))  # many scalar targets
>>> pl.Series("lists", [[0, 1], [0, 2], [1, 4]]).search_sorted(pl.lit([0, 2]))  # one list-typed value

Raise an error on struct.rename_fields() with an incorrect number of fields

Passing a number of names that doesn't match the number of struct fields now raises a SchemaError, instead of silently truncating.

Example

Before:

>>> s = pl.Series("s", [{"a": 1, "b": 2}])
>>> s.struct.rename_fields(["x"])
shape: (1,)
Series: 's' [struct[1]]
[
        {1}
]

After:

>>> s.struct.rename_fields(["x"])
Traceback (most recent call last):
...
polars.exceptions.SchemaError: struct.rename_fields() argument has a different number of fields than the struct it operates on (1 vs 2).

Hint: use struct.drop() to drop fields from the struct first.

Use instead, to drop fields first:

>>> s.struct.drop(["b"]).struct.rename_fields(["x"])

Remove the ordering parameter and the 'physical'/'lexical' special case on pl.Categorical()

Categoricals are now always ordered lexically, so the ordering parameter has been removed. Additionally, the special case where pl.Categorical("physical") / pl.Categorical("lexical") meant an ordering (rather than a category-pool name) has been removed.

Danger

pl.Categorical("physical") no longer means anything about ordering. It silently becomes a named Categories pool called "physical" instead of raising or warning.

Example

Before:

>>> pl.Categorical(ordering="lexical")
Categorical

After:

>>> pl.Categorical(ordering="lexical")
Traceback (most recent call last):
...
TypeError: Categorical.__init__() got an unexpected keyword argument 'ordering'

And the silent case:

>>> pl.Categorical("physical")
Categorical(Categories("physical"))

CSV reading changes

Set csv infer_schema_files default value to 10

Multi-file CSV scans now only inspect the first 10 files (instead of all of them) to infer the schema by default.

Example

Before:

>>> pl.scan_csv([b"a\n1"] * 10 + [b"a\nA"]).collect().tail(1)
shape: (1, 1)
┌─────┐
│ a   │
│ --- │
│ str │
╞═════╡
│ A   │
└─────┘

After:

>>> pl.scan_csv([b"a\n1"] * 10 + [b"a\nA"]).collect()
Traceback (most recent call last):
...
polars.exceptions.ComputeError: could not parse `A` as dtype `i64` at column 'a' (column number 1)

The current offset in the file is 0 bytes.

You might want to try:
- increasing `infer_schema_length` (e.g. `infer_schema_length=10000`),
- specifying correct dtype with the `schema_overrides` argument
- setting `ignore_errors` to `True`,
- adding `A` to the `null_values` list.

Use instead:

>>> pl.scan_csv([b"a\n1"] * 10 + [b"a\nA"], infer_schema_files=11).collect()

Start csv column name counting at 0

Auto-generated column names for headerless files now start at column_0 instead of column_1. This also affects read_excel/read_ods with has_header=False.

Example

Before:

>>> pl.read_csv(b"a,1,10;b,2,20", eol_char=";", has_header=False).columns
['column_1', 'column_2', 'column_3']

After:

>>> pl.read_csv(b"a,1,10;b,2,20", eol_char=";", has_header=False).columns
['column_0', 'column_1', 'column_2']

Respect file column order when schema is given for scan_csv()

A user-provided schema for read_csv/scan_csv is now matched by column name against the file's header, and the file's column order is respected, rather than matching positionally. New extra_columns/missing_columns parameters control name mismatches; their literal default is None, which behaves as "raise".

Example

Before:

>>> buf = b"a,b\nA,B\n"
>>> pl.scan_csv(buf, schema={"b": pl.String, "a": pl.String}).collect()
shape: (1, 2)
┌─────┬─────┐
│ b   ┆ a   │
│ --- ┆ --- │
│ str ┆ str │
╞═════╪═════╡
│ A   ┆ B   │
└─────┴─────┘

After:

>>> pl.scan_csv(buf, schema={"b": pl.String, "a": pl.String}).collect()
shape: (1, 2)
┌─────┬─────┐
│ b   ┆ a   │
│ --- ┆ --- │
│ str ┆ str │
╞═════╪═════╡
│ B   ┆ A   │
└─────┴─────┘

Note the values: previously schema was matched positionally against the file (so the first schema entry, b, silently received the file's first column, a's data). Now b correctly receives the file's actual b column.

Use instead, if the file has extra or missing columns relative to the schema:

>>> pl.scan_csv(buf, schema={"b": pl.String}, extra_columns="ignore").collect()
>>> pl.scan_csv(buf, schema={"a": pl.String, "b": pl.String, "c": pl.Int64}, missing_columns="insert").collect()

Automatically disable raise_if_empty when has_header=False and schema is given

raise_if_empty (on read_csv/scan_csv) now has a literal default of None, resolved internally as schema is None or has_header. In the narrow case of an empty, headerless source read with an explicit schema, this returns an empty frame with that schema instead of raising NoDataError. Pass raise_if_empty=True explicitly to keep the old behavior.

A related, undocumented default change: truncate_ragged_lines moves from a literal False to None, resolved as extra_columns == "ignore". The effective default is unchanged unless you also pass extra_columns="ignore", which now implicitly enables ragged-line truncation too. Passing them in direct conflict (extra_columns="ignore" together with an explicit truncate_ragged_lines=False) is a new hard error rather than the extra columns silently winning or losing:

>>> pl.read_csv(b"a,b\n1,2,3\n", extra_columns="ignore", truncate_ragged_lines=False)
Traceback (most recent call last):
...
ValueError: cannot set truncate_ragged_lines=False with extra_columns='ignore'
>>> pl.read_csv(b"", has_header=False, schema={"a": pl.Int64})
shape: (0, 1)
┌─────┐
│ a   │
│ --- │
│ i64 │
╞═════╡
└─────┘

Removal of deprecated functionality

Attempting to use most of the removed functionality below now raises one of two new exception types, each with a hint towards the replacement: polars.exceptions.AttributeRemovedError (an AttributeError) for removed attributes and methods, and polars.exceptions.ArgumentRemovedError (a TypeError) for removed/renamed parameters. Coverage is not complete, though: the entries marked below with [^1] still raise a plain built-in AttributeError or TypeError from ordinary attribute or argument resolution rather than the new typed exceptions.

Removed Deprecated in Use instead
DataFrame/LazyFrame.melt() 1.0 unpivot(), with index instead of id_vars and on instead of value_vars
DataFrame/LazyFrame.with_row_count() 0.20 with_row_index() (default column name is now "index", not "row_nr")
DataFrame/LazyFrame.approx_n_unique() 1.0 select(pl.all().approx_n_unique())
LazyFrame.fetch() 0.19 collect() in conjunction with head()
LazyFrame.profile() 1.43 none (misleading under the streaming engine)
DataFrame/LazyFrame.group_by(...).count() 0.20 len()
DataFrame.__dataframe__() 1.40 to_arrow()/to_pandas() (see above)
polars.interchange submodules and protocol symbols1 1.40 none (only CompatLevel remains; see above)
pl.Array.width 0.20 size
polars.testing.parametric.columns()1 0.20 column, in conjunction with a list comprehension
polars.testing.parametric.create_list_strategy()1 0.20 lists
Expr/Series.dt.datetime() 1.0 dt.replace_time_zone(None)
Expr/Series.dt.with_time_unit() 1.0 cast to Int64, then cast to the desired dtype
Expr.agg_groups(), pl.groups() 1.0 df.with_row_index().group_by(...).agg(pl.col("index"))
Expr.flatten() 1.0 list.explode(keep_nulls=False, empty_as_null=False) (note: different null semantics)
Expr.rechunk() 1.0 df.rechunk() after collecting
Series.cat.get_categories() 1.0 Series.unique(), or dtype.categories for an Enum
Series.cat.is_local(), to_local(), uses_lexical_ordering() 1.0 none (categoricals no longer have a local scope; ordering is always lexical)
Enum.union() / Enum.__or__ 1.38 construct explicitly: pl.Enum(pl.concat([e1.categories, e2.categories]).unique(maintain_order=True))
Config.set_auto_structify() 1.32 explicit pl.struct(...)
read_csv_batched() 1.37 scan_csv().collect_batches()
Expr/Series.str.explode()1 1.0 str.split("").explode() (note: empty strings now become null instead of being preserved)
read_csv/scan_csv(missing_utf8_is_empty_string=...)1 1.43 empty_string_is_null=... (note: the boolean is inverted)
pl.corr(ddof=...)1 1.17 none (parameter had no effect)
scan_ipc(memory_map=...)1 1.40 none (parameter had no effect)
read_parquet/scan_parquet(allow_missing_columns=...) 1.30 missing_columns="insert" / "raise"
read/scan_*(rechunk=...) 1.x .rechunk() on the result after reading
LazyFrame.map_batches(no_optimizations=...) 1.30 none (pushdown parameters already default to False)
DataFrame.to_arrow/write_ipc/write_ipc_stream(future=...) 1.1 compat_level=...
DataFrame.glimpse(return_as_string=...) 1.35 return_type="string"
DataFrame/LazyFrame.top_k/bottom_k(descending=...) 1.0 reverse=...
DataFrame/LazyFrame.rolling/group_by_dynamic(by=...), DataFrame.upsample(by=...) 0.20 group_by=...
DataFrame/LazyFrame.join(join_nulls=...) 1.24 nulls_equal=...
DataFrame.pivot(columns=...) 1.0 on=...
pl.Array(width=...) 0.20 shape=...
pl.from_dataframe(allow_copy=...) 1.23 none (removed along with the interchange protocol)
LazyFrame.collect/collect_async/explain/show_graph(streaming=...) 1.25 engine="streaming"
pl.collect_all/collect_all_async(streaming=...) 1.25 engine="streaming"
LazyFrame.explain(tree_format=...) 0.20 format="tree"
LazyFrame.unpivot(streamable=...) 1.5 none (parameter had no effect)
LazyFrame.sink_*(retries=...), DataFrame.write_*(retries=...) 1.37 max_retries in storage_options
The individual optimization-flag parameters (no_optimization, predicate_pushdown, ...)2 1.30 the optimizations parameter
LazyFrame.with_context() 1.0 pl.concat(..., how="horizontal")
pl.threadpool_size() 0.20.7 pl.thread_pool_size()
Expr.where() 0.20.4 filter()
Expr.shrink_dtype() 1.33 Series.shrink_dtype() (after collecting)
Expr.register_plugin() 0.20.16 polars.plugins.register_plugin_function()
Expr.from_json() 0.20.11 Expr.deserialize() (note: takes a file-like input, not a string)
Expr.meta.write_json() 0.20.11 Expr.meta.serialize()
Expr/Series.str.concat() 1.0 str.join() (note: the default delimiter is "", not "-")
Series.has_validity() 0.20.30 has_nulls()
Series.dt.mean(), Series.dt.median() 1.0 Series.mean(), Series.median()
polars.type_aliases module1 1.0 polars._typing
join(how="outer"), join(how="outer_coalesce")1 0.20.29 how="full" / how="full", coalesce=True (raises ValueError)
DataFrame/Series.to_numpy(use_pyarrow=...) 0.20.28 none (the native engine is used); .to_arrow().to_numpy() for PyArrow's
Series.to_numpy(zero_copy_only=...) 0.20.10 allow_copy=... (note: the boolean is inverted)
Expr/Series.replace(default=...), (return_dtype=...) 1.0 replace_strict()
Expr.map_batches(agg_list=...) 1.32 expr.implode().map_batches(...)
Expr.map_elements(returns_scalar=...) 1.32 none
Expr/Series.str.json_decode(infer_schema_length=...) 1.33 none
Expr/Series.rolling_*(min_periods=...), pl.rolling_cov/rolling_corr(min_periods=...) 1.21 min_samples=...
Series.equals(strict=...) 0.20.31 check_dtypes=...
assert_series_equal/assert_frame_equal(check_dtype=...) 0.20.31 check_dtypes=...
read_csv/scan_csv(dtypes=...) 0.20.31 schema_overrides=...
read/scan_csv/_ipc/_ndjson/_parquet(row_count_name=, row_count_offset=...) 0.20.4 row_index_name=... / row_index_offset=...
scan_ipc(cache=, file_cache_ttl=...) 1.40 none (the file cache is no longer supported)
read_ndjson/scan_ndjson(file_cache_ttl=...) 1.39 none (idem)
read_excel/read_ods(xlsx2csv_options=...) 0.20.6 engine_options=...
read_excel/read_ods(read_csv_options=...) 0.20.7 read_options=...
DataFrame.write_delta(overwrite_schema=...) 0.20.14 delta_write_options={"schema_mode": "overwrite"}
pl.from_repr(tbl=...) 0.20.17 data=...
pl.SQLContext(eager_execution=...) 0.20.31 eager=...
pl.QueryOptFlags(collapse_joins=...) 1.33.1 predicate_pushdown=...
polars.testing.parametric.series/dataframes(size=...) 1.0 min_size=... / max_size=...
Expr.list.to_struct(n_field_strategy=, upper_bound=...) 1.20 fields (see above)
  • The removed polars.interchange submodules and protocol symbols, the polars.testing.parametric strategy helpers, and the polars.type_aliases module raise ImportError/ModuleNotFoundError on import, and AttributeError on attribute access.
  • Expr/Series.str.explode() raises a plain AttributeError. Python's own suggestion machinery fires (Did you mean: 'encode'?) but points somewhere unhelpful.
  • read_csv/scan_csv(missing_utf8_is_empty_string=...), pl.corr(ddof=...) and scan_ipc(memory_map=...) raise Python's own TypeError for an unexpected keyword argument, with no pointer to the replacement.
  • join(how="outer")/how="outer_coalesce" raise a ValueError, but one that names the replacement.

Example

Before:

>>> lf.melt(id_vars="a", value_vars="b")
>>> df.with_row_count()
>>> df.join(df, on="a", join_nulls=True)

After:

>>> lf.melt(id_vars="a", value_vars="b")
Traceback (most recent call last):
...
polars.exceptions.AttributeRemovedError: `melt` was removed in version 2.0; use `LazyFrame.unpivot` instead, with `index` instead of `id_vars` and `on` instead of `value_vars`
>>> df.with_row_count()
Traceback (most recent call last):
...
polars.exceptions.AttributeRemovedError: `with_row_count` was removed in version 2.0; use `with_row_index` instead. Note that the default column name has changed from 'row_nr' to 'index'.
>>> df.join(df, on="a", join_nulls=True)
Traceback (most recent call last):
...
polars.exceptions.ArgumentRemovedError: the argument 'join_nulls' for 'DataFrame.join' was deprecated in version 1.24 and has been removed in 2.0. It was renamed to 'nulls_equal' in version 2.0.

Deprecations

Mark read_avro()/write_avro() as unstable

Avro support may change at any point without it being considered a breaking change. Enable pl.Config.warn_unstable(True) to be warned when using it.

Deprecate redundant strict=True on horizontal pl.concat()/pl.union()

Since how="horizontal" now always requires equal heights (see above), passing strict=True explicitly no longer does anything and is deprecated. Simply omit it.

Deprecate cut()/qcut() in favour of the bin_* methods

Expr/Series.cut() and Expr/Series.qcut() are deprecated in favour of bin_intervals(), bin_quantiles() and bin_ranks(). The new methods return an Enum of the labels you pass, or the integer bin index if you pass labels=False.

Deprecated Use instead
cut(breaks) bin_intervals(breaks, labels=..., right_closed=True)
qcut(quantiles) bin_quantiles(quantiles, labels=..., right_closed=True)
qcut(n) bin_quantiles(n, labels=..., right_closed=True)

bin_ranks() is a third option with no cut/qcut equivalent: it bins on position in sorted order, so the bins are near-equal-sized even when values repeat, whereas qcut() and bin_quantiles() keep equal values together and may produce empty (or imbalanced) bins.

Note the differences when translating:

  • The bin_* methods are left-closed by default. Pass right_closed=True to keep the right-closed bins of cut()/qcut(). (bin_ranks() is positional, so it has no right_closed.)
  • labels is required. There is currently no equivalent of the automatic interval-string labels ("(-inf, -1]" etc.) that cut()/qcut() generate by default.
  • include_breaks=True becomes include_intervals=True, which returns a struct with bin, left and right fields rather than breakpoint and category.
  • qcut()'s allow_duplicates has no counterpart: bin_quantiles() accepts duplicate quantiles and simply produces an empty bin between them.

Example

Before:

>>> df = pl.DataFrame({"foo": [-2, -1, 0, 1, 2]})
>>> df.with_columns(pl.col("foo").cut([-1, 1], labels=["a", "b", "c"]).alias("bin"))

After:

>>> df.with_columns(
...     pl.col("foo")
...     .bin_intervals([-1, 1], labels=["a", "b", "c"], right_closed=True)
...     .alias("bin")
... )

  1. These raise a plain built-in exception rather than one of the new typed ones, because the symbol or parameter is simply gone rather than routed through the removal machinery: ↩↩↩↩↩↩↩↩↩

  2. The full set: no_optimization, type_coercion, predicate_pushdown, projection_pushdown, simplify_expression, slice_pushdown, comm_subplan_elim, comm_subexpr_elim, cluster_with_columns, collapse_joins, plus the private _eager, _type_check, and _check_order, on LazyFrame.explain/show_graph/collect and pl.collect_all. Pass a pl.QueryOptFlags object via optimizations= instead, e.g. lf.collect(optimizations=pl.QueryOptFlags(predicate_pushdown=False)). ↩