
Data Expert
A local-first workbench for data engineers. Point it at a folder, open any Parquet, CSV, JSON or Excel file, and browse, search, profile, edit and reshape it in your browser. Every save becomes a version you can compare and restore, and an AI assistant turns plain-language requests into transformations — without ever sending your data to the model.
FastAPI · Python · Polars · PyArrow · React 18 · TypeScript · Vite · Tailwind · Open Source
Background
Data engineers need a fast, local tool to explore and transform large datasets without uploading sensitive files to cloud services or juggling separate utilities.
Responsibilities
Build a local-first workbench that browses, searches, profiles, edits and versions datasets from kilobytes to gigabytes, with an AI assistant that can drive transformations.
Achievements
- Implemented a FastAPI backend with lazy Polars/PyArrow reads so multi-gigabyte Parquet files open on metadata and page rows on demand.
- Added a compact background search index so large files can be searched without scanning the whole file each time.
- Shipped safe editing with pending changes, undo, atomic saves, and git-like version history stored as kilobyte-scale reverse diffs.
- Built a privacy-first AI assistant that plans steps from a function library, writes sandboxed functions, and runs locally — only schema, stats and truncated samples reach the model.
- Supported 14+ LLM providers including Ollama, LM Studio, vLLM, OpenAI, Anthropic, Gemini and OpenRouter.
- Added profiling, export/combine tooling, and a React 18 + TypeScript + Vite + Tailwind UI.
Outcomes
A self-contained, MIT-licensed data workbench that opens huge files instantly, versions every edit, and lets an AI plan and run transformations entirely on the user's machine.