S3 & URL

View and query Parquet on S3 in your browser

Checking one Parquet file in a bucket should not require aws s3 cp, a notebook, or an Athena query. viewparquet points DuckDB-WASM straight at the object: paste a public HTTPS or presigned URL, an s3:// path, or connect a private bucket with access keys.

For a Parquet endpoint that supports HTTP range requests, DuckDB can request portions of the object instead of downloading the entire file up front. Actual transfer depends on the query, Parquet layout, and host behavior. Requests go from your browser to the object store; private keys and session tokens are not persisted, while presigned URLs are kept only in session storage because their query strings contain temporary authorization. AWS S3 and compatible endpoints such as R2, GCS with HMAC interoperability, and MinIO require the appropriate endpoint, credentials, permissions, HTTPS setup, and CORS configuration.

Updated August 21, 2026 · Reproducible checks, limitations, and primary references are included below.

Open remote Parquet

Reproduce a remote read, then connect your bucket

  1. Reproduce the public-URL path

    Open the remote dialog, choose Public or presigned, and paste https://viewparquet.com/data/flights-200k.parquet with Auto-detect selected. Run SELECT count(*) AS rows FROM data; the repository copy contains 231,083 rows. This confirms the dialog and query path, not that every third-party host permits browser reads or byte ranges.

  2. Choose the least-secret connection

    Use a public HTTPS URL when the object is public, or paste an unmodified presigned URL when temporary access is enough. For a private bucket, choose Private bucket and enter an s3://bucket/key path plus access keys or temporary STS credentials. Never place credentials in an ordinary URL.

  3. Match the provider settings

    AWS needs the bucket region. R2 uses <account_id>.r2.cloudflarestorage.com, region auto, and path-style addressing. GCS uses an HMAC interoperability key with storage.googleapis.com. MinIO/custom uses its host and port, usually path-style, with HTTPS required from the hosted app unless the browser permits a local-development exception.

  4. Verify CORS and range behavior

    In DevTools → Network, the object request should go directly to your storage host. A successful range response is normally 206 Partial Content with Content-Range; a 200 response can mean the host ignored the range and may transfer more data. CORS must allow this site's origin, GET and HEAD, request headers used by signing, and readable range-related response headers.

  5. Query narrowly, then export

    Start with a row count or a small projection before a full scan. Predicate and column pushdown can reduce reads only when the query, statistics, row-group layout, and endpoint permit it. Exports are generated in the browser, but creating one can still require reading every selected value.

What actually proves a selective remote read

“Opened from S3” and “downloaded only the needed bytes” are different claims. Use the browser network evidence and the query shape together before describing the read as selective.

Confirm the request destination

The object request should go from the browser to the configured storage host. A proxy, CDN, or signed URL may change the observable host and can also change range behavior.

Confirm that a range was honored

Inspect the request and response: a Range request answered with 206 Partial Content and Content-Range is direct evidence for that byte range. A 200 response does not prove a selective transfer.

Compare a narrow query with a full scan

Select one column with a useful predicate, note transferred bytes, then compare with SELECT * or an export. Column and predicate pushdown help only when the file statistics, row groups, query, and endpoint allow them.

Start with a controlled object before debugging a bucket

Load https://viewparquet.com/data/flights-200k.parquet through the remote dialog and run this query while the Network panel is open.

SELECT count(*) AS rows
FROM data
WHERE distance >= 1500;

What to expect: The query should run against the repository sample, but the transferred-byte result depends on the browser build and request sequence. Record the response status, Content-Range when present, and bytes transferred rather than turning one run into a universal performance claim.

Primary references

These sources support the format and tool behavior described on this page. Product-specific boundaries are checked against the shipped viewer using the claim-verification method.

Common questions

Can I view a Parquet file on S3 without downloading the whole file?

Only when the endpoint honors HTTP range requests and the query can use selective reads. DuckDB can request Parquet metadata and relevant portions of the object, but a host that returns 200 instead of 206, a full-table query, unsuitable row groups, missing statistics, or an export can still transfer most or all of the object.

How do I open a private S3 bucket safely in a browser viewer?

In viewparquet you connect with access keys or temporary STS tokens held in the current browser session and not persisted with the recent-source record. Requests go directly from the browser to S3; a presigned URL is an alternative that does not require entering keys in the viewer.

Which storage providers work besides AWS S3?

viewparquet has presets for Cloudflare R2, Google Cloud Storage through HMAC interoperability, and MinIO/custom S3-compatible endpoints, plus plain HTTPS and presigned URLs. R2 and GCS use their documented interoperability endpoints and path-style defaults; success still depends on credentials, permissions, region, addressing mode, TLS, CORS, and range support.

Why does my bucket file fail to load in the browser?

Start in the browser Network panel: no readable response usually points to CORS, 403 to permissions, signing, region, or expired authorization, 404 to the bucket/key, and 416 or an ignored Range header to endpoint behavior. The storage host must allow this origin, GET and HEAD, signing or Range request headers, and expose ETag, Content-Length, Content-Range, and Accept-Ranges when it sends them.

Why did a presigned Parquet URL stop working?

Presigned URLs expire and their signature covers specific request details, so use a newly generated URL and paste its complete, unmodified query string. A URL can also fail when the signing region or method is wrong, the object moved, the browser is blocked by CORS, or a proxy rewrites the URL; treat it as a bearer credential and clear Recent sources on a shared browser.

What CORS settings does a browser Parquet reader need?

Allow https://viewparquet.com as an origin, allow GET and HEAD, permit the Range and signing headers actually shown by the browser preflight, and expose ETag, Content-Length, Content-Range, and Accept-Ranges when available. Provider syntax differs, so begin with the tested example in the S3 guide and tighten allowed headers after inspecting your own OPTIONS request.

Do I still get privacy when reading from S3?

The object request goes from storage to your browser and SQL runs in DuckDB-WASM; viewparquet does not proxy the dataset, and analytics excludes remote URLs, bucket paths, credentials, SQL text, values, and exports. Optional AI is separate: invoking it sends your message and configured context directly to the provider you chose, and that context can include bounded result values when enabled.

Continue with a related task

Try it on your own file

Drop a Parquet, GeoParquet, CSV, or JSON file into viewparquet: browse rows, run DuckDB SQL, and export results in-browser without uploading the dataset to viewparquet. Optional AI sends messages and applicable context directly to the provider you choose; review Settings → AI before use.