NEW Qlik to dbt migration AI Proof Pricing Book a demo Get the free assessment

Hard sources

Also parsed

Targets

Warehouses

Runtimes

Before migration

After migration

Convert SAS programs to Polars

Enterprise Guide projects and DATA step parsed structurally. Emitted as Polars you run on a machine or Kubernetes, with DuckDB when PROC SQL is relational, and Iceberg when the lake is already the source.

Architecture

SAS in. Polars out, with the right engine for the job.

Deterministic parsers read the SAS estate and emit native Polars. Simple SQL stays in Polars. Heavy joins run in DuckDB on the same machine. Iceberg tables feed expressions without standing up Spark for a mid-size job.

SAS programs → SAS2PY parser → Polars, DuckDB, Iceberg

SAS
Base SAS DATA step / macros
DI Studio Jobs + mappings
EG / EM Projects + flows
Viya / CAS CASL + actions
SAS2PY Parser
Deterministic parse AI where it helps
Expression emit DATA step → LazyFrame
SQL routing Polars or DuckDB
Lake ingest Iceberg → Polars
Run
Polars LazyFrames + SQL
DuckDB Same machine
Iceberg Your catalog
Kubernetes When you need a fleet

MigryX AI handles the logic parsers cannot resolve alone, and every change it makes goes through the same parity checks. It runs on a model you approve, air-gapped if your estate requires it. DuckDB and Iceberg use what you already installed; we do not stand up a new warehouse.

Why Polars

Most SAS jobs are not lakehouse-scale

Spark tax on mid-size work

500MB to 40GB jobs pay a cluster bill they do not earn. Polars runs them in-process, lazily. DuckDB sits beside it when PROC SQL needs real joins.

EG projects are not Spark jobs

The parser reads the flow and emits a LazyFrame plan, not a notebook full of collect().

The lake is already there

If the heavy tables live in Iceberg, we scan them into Polars, or let DuckDB query Iceberg and hand the result back. Same catalog you already configured.

You still need parity

Same row-level compare. Different engine.

How the job runs

Four modes. One Polars program.

DATA step always becomes Polars expressions. PROC SQL and lake reads pick the engine that fits the workload, chosen per conversion, not hardcoded in the SAS.

Polars SQL

Default for simple SELECT / filter / project on in-memory frames.

Short PROC SQL stays in Polars. No extra process. Fits filters, projections, and light aggregations.

DuckDB SQL

When PROC SQL looks like Spark SQL: joins, CTEs, windows.

DuckDB runs on the same machine as Polars, registers the frames, executes the SQL, and returns a Polars DataFrame. Closest in-process stand-in for a Spark SQL job without a cluster.

Iceberg → expressions

Little SQL. A lot of lake data. DATA step still wins.

Scan Iceberg from the catalog you already run. Then filter, derive, and aggregate in Polars, the same LazyFrame path as a SAS dataset.

DuckDB on Iceberg

Relational SQL over lake tables, result back in Polars.

DuckDB reads Iceberg directly. We do not pull the whole table into Polars first. Warehouse pass-through (Oracle, Teradata, …) is unchanged, those connections still go to the database.

DuckDB is assumed installed next to Polars. Iceberg credentials stay in your existing catalog setup.

Parser output

SAS filter to a LazyFrame

A DATA step subset plus PROC MEANS, emitted as a lazy Polars plan that does not run until sink.

SAS
/* SAS */
data gold;
  set txn;
  if amount > 1000;
run;
proc means data=gold noprint;
  class segment;
  var amount;
  output out=sum sum=;
run;
→
SAS2PY
converts
Polars
# DATA step + MEANS → Polars
import polars as pl
gold = pl.scan_parquet("txn.parquet").filter(
    pl.col("amount") > 1000
)
summary = gold.group_by("segment").agg(
    pl.col("amount").sum()
).collect()

scan_ + filter + group_by is the plan. collect() is the only action.

Parser output

PROC SQL that Spark would have handled

Polars SQL is the simple path. Joins and windows go to DuckDB on the same box, then back to Polars.

SAS
/* SAS */
proc sql;
  create table work.out as
  select a.segment,
         sum(b.amount) as total
  from gold a
  inner join lookup b
    on a.id = b.id
  group by a.segment;
quit;
→
SAS2PY
converts
DuckDB → Polars
# Relational SQL on the same machine
import duckdb
out = duckdb.sql("""
  SELECT a.segment, sum(b.amount) AS total
  FROM gold a
  JOIN lookup b ON a.id = b.id
  GROUP BY a.segment
""").pl()

Iceberg sources use your catalog: scan into Polars for DATA-step work, or query in DuckDB when the SQL is the point.

Coverage

SAS to Polars: artifact mapping

SASPolars pathNotes
DATA stepLazyFrame / expressionsLazy until sink
Simple PROC SQLPolars SQLSelect, filter, light agg
Relational PROC SQLDuckDB, result as PolarsJoins, CTEs, windows
Large Iceberg, little SQLIceberg scan → expressionsExisting catalog
SQL over IcebergDuckDB reads IcebergNo full-table preload
EG projectPython moduleOne flow, one file
SAS datasetParquet / IPC / IcebergColumnar
Validation

How is the conversion proven? Every conversion validated to row-level parity

SAS output compared to Polars output: row by row, column by column. Differences flagged before sign-off. DuckDB and Iceberg paths use the same compare.

See how Data Matching works →
Guide

More on SAS to Polars

Construct mappings, code examples and platform notes for teams planning this move.

Migrating SAS to Polars with MigryX: LazyFrame Pipelines at Scale

SAS has been the backbone of enterprise analytics for decades, powering regulatory reporting, risk modeling, clinical trials, and data preparation across banking, insurance, healthcare, and government. But the economics and technical landscape have shifted. SAS licensing costs continue to climb, the talent pool is shrinking, and modern alternatives deliver superior performance at a fraction of the cost. The question is no longer whether to migrate, it is where to migrate to.

For organizations that need speed, simplicity, and Python-native tooling without the overhead of a Spark cluster, Polars is emerging as a compelling target. This article explains how MigryX converts SAS programs to idiomatic Polars LazyFrame pipelines, with practical examples and performance comparisons.

Why SAS to Polars?

Three forces are driving SAS-to-Polars migration decisions simultaneously.

Licensing costs. SAS licensing is notoriously expensive: annual costs for enterprise deployments routinely reach seven figures. Polars is MIT-licensed and free. The savings are not marginal; they are transformational, often funding the entire migration project and still leaving budget to spare.

The Python ecosystem. Python has become the lingua franca of data engineering, data science, and machine learning. Moving from SAS to Python means access to thousands of libraries, a massive talent pool, seamless integration with cloud platforms, and the ability to embed analytics into applications rather than running them in isolated SAS environments.

Polars' performance advantage. The traditional migration path from SAS has been to pandas, but pandas struggles with the dataset sizes that SAS handles routinely. A SAS program processing 50 million rows runs comfortably; the same logic in pandas can exhaust memory or take hours. Polars closes this gap. Its multi-threaded Rust engine and lazy evaluation deliver performance that meets or exceeds SAS for data preparation workloads, typically at 5-20x the speed for equivalent operations.

SAS to Polars Mapping

SAS constructs do not map one-to-one to Polars. Each requires careful decomposition, and the complexity varies significantly depending on the patterns involved.

SAS ConstructComplexityKey Challenge
DATA stepHighImplicit output, variable retention, and conditional logic require decomposition into LazyFrame chains
PROC SQLMediumSAS SQL extensions (calculated, monotonic, INTO) have no direct Polars equivalents
SAS MacrosHighText substitution semantics, nested macro calls, and conditional compilation require full expansion before translation
MERGE with BYMedium-HighSAS merge behavior differs from standard joins: especially with duplicate keys and IN= dataset options
RETAIN / BY-group stateHighRow-level state tracking across groups requires mapping to window expressions and cumulative functions

MigryX handles the full SAS construct landscape, generating idiomatic Polars code that leverages lazy evaluation and expression-based APIs.

Code Comparison: SAS DATA Step vs. Polars LazyFrame

Consider a typical SAS DATA step that merges two datasets, applies a filter, and computes a derived column. This pattern appears in virtually every SAS codebase.

SAS

data work.combined;
    merge work.orders(in=a) work.customers(in=b);
    by customer_id;
    if a and b;
    if order_amount > 100;
    profit = order_amount - cost;
    margin = profit / order_amount;
    length tier $10;
    if margin >= 0.3 then tier = 'HIGH';
    else if margin >= 0.15 then tier = 'MEDIUM';
    else tier = 'LOW';
run;

MigryX converts this multi-step DATA step (with its merge, filter, derived columns, and conditional logic) into an optimized Polars LazyFrame pipeline that leverages expressions and lazy evaluation for maximum performance. The Polars optimizer pushes filters down to the data source, skips reading unused columns, and executes joins and computations in parallel across all CPU cores. The SAS version processes rows sequentially in a single thread.

Handling SAS-Specific Patterns

SAS has idioms that do not map directly to standard DataFrame operations. MigryX handles these patterns with purpose-built translation logic.

RETAIN and Running State

SAS RETAIN statements (used for running totals, lag values, and group boundary detection) are among the trickiest patterns to translate. Polars' expression API handles these elegantly, but the translation requires deep understanding of both paradigms. MigryX handles all RETAIN patterns automatically.

BY-Group Processing

SAS DATA steps with BY statements process data group-by-group, with automatic FIRST. and LAST. variables. Polars handles this through .group_by() with aggregation expressions, or through .over() window expressions when row-level output is needed. MigryX detects BY-group patterns in the DATA step logic and selects the appropriate Polars construct.

ARRAY Processing

SAS arrays iterate over a set of columns within a DATA step: typically for scoring, recoding, or applying the same transformation to multiple variables:

array scores{5} score1-score5;
do i = 1 to 5;
    scores{i} = scores{i} * 1.1;
end;

SAS ARRAY processing requires decomposition into Polars horizontal operations: a non-trivial restructuring that MigryX handles automatically.

FORMAT and INFORMAT

SAS formats and informats control how values are displayed and read: a concept with no direct parallel in Python DataFrames. Date formats, numeric widths, and character informats all require careful translation to appropriate data types. MigryX automatically maps SAS formats and informats to appropriate Polars data types.

Get the free assessment on a sample of your code →

18
SAS engagements
16
on Spark

Proven with regulated enterprises

28 regulated enterprises, including six global systemically important banks, have modernized with MigryX. Customer names are shared under NDA in a demo, with reference calls on request.

See all engagements →