Columnar files New
Connect Parquet Folder to Excel, Sheets and AI
Query a folder of Parquet files as one table. Parquet carries its own column types, so nothing here is guessed — a decimal stays a decimal and a timestamp stays a timestamp.
One connection, every surface
Where your Parquet Folder data can go
Connect Parquet Folder once and the same read-only connection feeds all of these — no second setup, no second copy of the data.
Parquet Folder to Excel
Microsoft Excel · Excel add-in
Pull live Parquet Folder results straight into a worksheet and refresh them on demand — desktop Excel, Excel Online, Microsoft 365.
How Excel works no Parquet Folder walkthrough written yetParquet Folder to Google Sheets
Sheets add-on
Run a saved Parquet Folder query from the sidebar and drop the rows into the sheet. Shared collaborators can refresh it themselves.
How Google Sheets works no Parquet Folder walkthrough written yetParquet Folder MCP server
Claude, Cursor and MCP clients
Give an AI assistant read-only access to Parquet Folder with the schema it needs to write correct SQL — no credentials in the chat.
How MCP works no Parquet Folder walkthrough written yetParquet Folder REST API
HTTP endpoint
Publish a Parquet Folder query as an authenticated JSON endpoint any application can call, with an OpenAPI 3.1 spec and ready-made Postman, Insomnia and Hoppscotch collections. No database port is opened.
How REST API works no Parquet Folder walkthrough written yetParquet Folder to Airtable
Automation platform
Sync Parquet Folder rows into an Airtable base on a schedule, or fetch them inside an Airtable automation script.
How Airtable works no Parquet Folder walkthrough written yetParquet Folder to Baserow
Automation platform
Feed a Baserow table from Parquet Folder over the REST endpoint — self-hosted or Baserow cloud.
How Baserow works no Parquet Folder walkthrough written yetParquet Folder to SeaTable
Automation platform
Keep a SeaTable base current with Parquet Folder data without exporting a file or exposing the database.
How SeaTable works no Parquet Folder walkthrough written yetParquet Folder to Smartsheet
Automation platform
Push Parquet Folder results into a Smartsheet grid so plans and reports read from the source system, not last week's export.
How Smartsheet works no Parquet Folder walkthrough written yetParquet Folder to Anvil
Anvil Works · App platform
Back an Anvil Python app with Parquet Folder through the REST endpoint instead of embedding database credentials in the app.
How Anvil works no Parquet Folder walkthrough written yetParquet Folder to Power BI
Power Query M
Paste the generated Power Query M into the Power BI Advanced Editor and the report reads live Parquet Folder results over HTTPS — no ODBC driver, no database port opened.
How Power BI works no Parquet Folder walkthrough written yetParquet Folder alerts and reports
Slack · Discord · Email · Webhook
Put a Parquet Folder query on a schedule and have the rows delivered to Slack, Discord, email or a signed webhook — or hold the message until a row count, threshold or percentage change crosses the line you set.
How alerts and reports work no Parquet Folder walkthrough written yetHow it works
5 steps, no inbound firewall change
Install the Network Agent where it can see the folder — an export directory, a small data lake, a vendor drop.
Point the connector at the folder or the whole tree; every .parquet file is matched.
Parquet stores its schema inside each file, so the agent reads those descriptions rather than sampling rows. No type is inferred from text, and no sampling error is possible.
The files load into one table, stacked, with a _source_file column on every row naming its file.
Query it with the same SQL as your databases — and join it to them in the same statement.
Feature deep-dive
What Parquet Folder gives you
Why this one is exact
- Types come from the file itself — decimals, timestamps and integers arrive as what they already are.
- Column union — if some files carry a column the others lack, it simply reads as empty for those files instead of breaking the set.
- If two files genuinely disagree on a column's type, the connector refuses to pin rather than flattening both to text. You are told which file and which column.
-- Rows and revenue per export file
SELECT _source_file, COUNT(*) AS rows, SUM(amount) AS total
FROM data
GROUP BY _source_file
After the first sync
A later file that adds a column, drops a pinned one, or changes a type is held back and recorded in files_events by name. The rest of the folder keeps loading, so one malformed export never takes the table offline. Reading is cheap: only files that actually changed are re-read.
Shared by the whole File Set family
Every folder connector also gives you
- files_current — a live inventory: every file the connector can see right now, with its path, size and modified date.
- files_events — the audit trail: what appeared, what changed, what vanished, and anything held back, with the filename and the reason.
- directories and volumes — per-folder totals, daily growth history, and how much room is left on the drive.
- Only files that actually changed are re-read on each sync, so a folder of 50,000 files is not re-parsed because one new export landed.
Good to know
- Read-only, enforced — one statement at a time, SELECT and friends only. Your files are never written, moved or renamed.
- Nothing is uploaded — the data is cached, encrypted, on your own machine beside the agent. Only the result of a query leaves your network.
- Sensible defaults — up to 250,000 files, 16 folders deep, 512 MB per file, all adjustable. Recycle bins, .git and node_modules are always skipped.
- The SQL dialect is DuckDB — the same SQL you would write against any other connector.
- Requires Network Agent 2.6 or newer.
Cross-source SQL
Join Parquet Folder to the rest of your data
A folder of files is a set of SQL tables like any other, so one statement can join it to a database and an API at once. Each source runs only the part it can, streams the result back, and the join happens centrally — the sources never talk to each other and nothing is copied anywhere.
3 connections · 3 agents
One statement
-- nothing copied, nothing merged, nothing scheduled
SELECT c.region, COUNT(*) AS orders, SUM(i.amount_due) AS invoiced
FROM parquet_lake.fileset.data1 f
JOIN pg_crm.public.customers2 c ON c.id = f.customer_id
JOIN billing.stripe.invoices3 i ON i.customer = c.stripe_id
GROUP BY c.region
ORDER BY invoiced DESC;
The three parts are connection, schema and table — and the connection name is whatever you called it. Illustrative columns; your tables will be your tables. Read-only applies to every piece: SELECT, WITH and EXPLAIN only, with a ceiling on how much any one source may hand over for a single query. How federated queries work
Connection details
What Parquet Folder needs
- Folder
- One or more roots — local disk, mapped drive or UNC share; every .parquet file is matched, subfolders included
- Schema
- Read from each file's own Parquet footer — no sampling, no inference, no type worked out from text
- Column union
- A column only some files carry reads as empty in the others rather than breaking the set
- Type conflicts
- If two files genuinely disagree on a column's type the connector refuses to pin and names the file and the column, instead of flattening both to text
- Tables
- One table, every file stacked into it, each row carrying a _source_file column naming its file
- Limits
- Up to 250,000 files, 16 folder levels deep, 512 MB per file — all adjustable
- Credentials
- None — there is no server; the agent reads the files in place
- SQL dialect
- DuckDB — standard SQL, nothing folder-specific to learn
- Agent
- Network Agent 2.6 or newer
There is no database server in this picture. The agent reads the .parquet files where they already live, caches the rows in an encrypted DuckDB store on the same machine, and re-reads only files that actually changed — a folder of 50,000 files is not re-parsed because one new export landed. Nothing is uploaded to Query Streams; the only thing that ever leaves your network is the result of a query.
Parquet is the one format in this family where nothing has to be inferred. Every other folder connector works out what a column is by looking at values; Parquet already knows, so the agent reads the description rather than the data. That makes setup cheap even on very large folders — only the footers are examined — and it removes the whole class of problem where a column looks numeric for the first thousand rows and then turns out not to be.
Vendor documentation: parquet.apache.org
FAQ
Questions about Parquet Folder
Which tools can read Parquet Folder data through Query Streams?
All of them, from one connection: Excel, Google Sheets, MCP, REST API, Airtable, Baserow, SeaTable, Smartsheet, Anvil, Power BI, scheduled alerts and reports. Connect the folder once and every surface reads the same read-only connection — there is no per-tool setup and no second copy of the data.
Do my files get uploaded to Query Streams?
No. The Network Agent reads the files in place and caches rows in an encrypted store on the same machine. The files themselves never leave your network — only the result rows of a query do, over a single outbound encrypted connection with no inbound firewall port.
Can Query Streams change, move or rename my files?
No. Files are opened strictly read-only and are never written, moved or renamed. Queries are enforced read-only at the point of execution — one statement at a time, SELECT and friends only.
What does Query Streams need to connect to Parquet Folder?
A folder path the agent machine can see — no server, no credentials, no drivers to install. Folder: One or more roots — local disk, mapped drive or UNC share; every .parquet file is matched, subfolders included. Schema: Read from each file's own Parquet footer — no sampling, no inference, no type worked out from text. Column union: A column only some files carry reads as empty in the others rather than breaking the set. Type conflicts: If two files genuinely disagree on a column's type the connector refuses to pin and names the file and the column, instead of flattening both to text.
Can I join a folder of files to a database in the same query?
Yes — that is a federated query. One statement can reference Parquet Folder and your other connections at once, written as connection.schema.table. Each source runs only the part it can and streams the result back; the join happens centrally, so the sources never connect to each other and nothing is copied or scheduled. Read-only applies to every piece — SELECT, WITH and EXPLAIN only — and there is a ceiling on how much any one source may hand over for a single query. Federated queries are a plan feature; the federated queries page carries the current source and size limits.
Put Parquet Folder where the work happens
Install the agent, point it at your folder, and pick a destination.
Read-only Outbound only Credentials stay on the agent

