Skip to main content

Pandas

SigilYX can read and write YXDB files as Pandas DataFrames. Data flows through PyArrow under the hood, so you need the [pandas] extra installed.

Installation

pip install sigilyx[pandas]

Reading

import sigilyx as yx

df = yx.read_yxdb_pandas("data.yxdb")
print(type(df)) # <class 'pandas.core.frame.DataFrame'>

The read path is: Rust reader -> Arrow arrays -> PyArrow Table -> Pandas DataFrame. Despite the conversion steps, this is still orders of magnitude faster than pure-Python YXDB readers.

Writing

import sigilyx as yx
import pandas as pd

df = pd.read_csv("data.csv")
yx.write_yxdb_pandas("output.yxdb", df)

Pandas DataFrames are converted to PyArrow Tables before being passed to the Rust writer.

Type Mapping

Pandas dtypeYXDB TypeNotes
boolBoolean
int16Int16
int32Int32
int64Int64
float32Float
float64Double
object (strings)V_WStringVariable-length UTF-16
datetime64[ns]DateTimeConverted to microseconds
object (date)Date
object (time)Timedatetime.time values
object (Decimal)FixedDecimaldecimal.Decimal values
bytesBlobRaw binary data
Prefer Polars for performance

The Polars path (read_yxdb()) is the fastest because it avoids the Pandas conversion overhead. If you're starting a new project and don't have an existing Pandas dependency, consider using Polars directly. See the Polars guide.

Common Patterns

Read YXDB, process with Pandas, write back

import sigilyx as yx

df = yx.read_yxdb_pandas("input.yxdb")

# Standard Pandas operations
df["total"] = df["quantity"] * df["price"]
df = df[df["total"] > 100]
df = df.sort_values("total", ascending=False)

yx.write_yxdb_pandas("output.yxdb", df)

Convert YXDB to CSV

import sigilyx as yx

df = yx.read_yxdb_pandas("data.yxdb")
df.to_csv("data.csv", index=False)

Convert Excel to YXDB

import sigilyx as yx
import pandas as pd

df = pd.read_excel("data.xlsx")
yx.write_yxdb_pandas("data.yxdb", df)

Interop with scikit-learn

import sigilyx as yx
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

df = yx.read_yxdb_pandas("training_data.yxdb")

X = df[["feature_1", "feature_2", "feature_3"]]
y = df["target"]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
model = LinearRegression().fit(X_train, y_train)
print(f"R² = {model.score(X_test, y_test):.3f}")