Turkish LLM Preference & Alignment Dataset
A curated Turkish-language preference dataset designed for LLM training, alignment, preference optimization, and evaluation.
The dataset contains approximately 600 chosen/rejected response pairs. Each example is structured to distinguish a higher-quality preferred response from a lower-quality alternative.
LLM preference learning
DPO / preference optimization
Reward model training
Turkish-language alignment
Model evaluation and benchmarking
Hard-negative and failure-mode analysis
JSONL with structured chosen/rejected response pairs.
Turkish
The dataset was deliberately curated rather than produced as a raw bulk generation set. Particular attention was given to response quality, difficult distinctions between competing answers, and examples useful for identifying model weaknesses.
Suitable for research, experimentation, evaluation, and commercial AI development subject to the applicable dataset license.
Learn how to interact with this dataset using the Ouro SDK or REST API.
API access requires an API key. Create one in Settings → API Keys, then set OURO_API_KEY in your environment.
Get dataset metadata including name, visibility, description, and other asset properties.
Get column definitions for the underlying table, including column names, data types, and constraints.
| Column | Type |
|---|---|
| chosen | text |
| id | uuid |
| prompt | text |
| rejected | text |
Fetch the dataset's rows. Use query() for smaller datasets or load() with the table name for faster access to large datasets.
Update dataset metadata (visibility, description, etc.) and optionally write new rows to the table. Writing new data will replace the existing data in the table. Requires write or admin permission on the dataset.
# Get column definitions for the underlying table
columns = ouro.datasets.schema(dataset_id)
for col in columns:
print(col["column_name"], col["data_type"]) # e.g., age integer, name text# Option 1: All rows as a Pandas DataFrame
df = ouro.datasets.query(dataset_id)
print(df.head())
# Option 2: Read-only SQL — pass a query string; use {{table}} as the placeholder
agg = ouro.datasets.query(
dataset_id,
"SELECT col, count(*) AS n FROM {{table}} GROUP BY col ORDER BY n DESC",
)import pandas as pd
# Update dataset metadata
updated = ouro.datasets.update(
dataset_id,
visibility="private",
description="Updated description"
)
# Update dataset data (replaces existing data)
data_update = pd.DataFrame([
{"name": "Charlie", "age": 33},
{"name": "Diana", "age": 28},
])
updated = ouro.datasets.update(dataset_id, data=data_update)import os
from ouro import Ouro
# Set OURO_API_KEY in your environment or replace os.environ.get("OURO_API_KEY")
ouro = Ouro(api_key=os.environ.get("OURO_API_KEY"))
dataset_id = "82b7fe9e-1f74-4fb6-ac6a-4fc6849dc5b6"
# Retrieve dataset metadata
dataset = ouro.datasets.retrieve(dataset_id)
print(dataset.name, dataset.visibility)
print(dataset.metadata)