Developer Tool · by Ralfo Becher, RALFORION

qvd2parquet

Convert Qlik QVD files to Parquet, exactly. Built for a Qlik to Dremio migration, not for a weekend.

Free and open source, Apache 2.0, a single Go binary with no runtime to install — github.com/ralforion/qvd2parquet

What it is

qvd2parquet is a command-line converter from Qlik QVD files to Apache Parquet. It was built for a real Qlik to Dremio migration, where a QVD layer was the last thing standing between the Qlik estate and the lakehouse, and where getting those files out correctly turned out to be harder than it looks.

qvd2parquet input.qvd output.parquet

It reads standard, unencrypted QVD files, preserves useful Parquet types instead of stringifying everything, keeps MONEY and FIX columns as exact Parquet decimals, decodes records in parallel while keeping the rows in the order the QVD holds them, and streams batches into Parquet row groups, so a large file never has to be materialized in memory.

The output is the deliverable, not the log line. A quality gate, on by default, reads the written Parquet back and compares it against metrics collected from the values the converter actually produced, before the file is renamed into place. A conversion that says it succeeded has been checked.

On a real SAP extract

Terminal session converting a 28 MB SAP BSEG QVD extract to Parquet: qvd2parquet --inspect resolves 15 columns in 16ms, then the conversion writes 2,400,000 rows and 15 columns to 5.9 MiB of Parquet in 2.025 seconds, with MONEY fields DMBTR and WRBTR typed as decimal(7,2) and zero-padded codes BELNR, HKONT and KOSTL kept as utf8.
A 28 MB SAP BSEG extract: 2.4 million rows across 15 columns. --inspect resolves the whole schema in 16ms by reading 429 KiB of symbol tables and skipping the 27.5 MiB record area entirely. The conversion itself writes 5.9 MiB of Parquet in 2.025 seconds. Note what the schema decided on its own: DMBTR and WRBTR are decimal(7,2) rather than doubles, BUDAT and CPUDT resolve to date32 with no declared type in the file, and BELNR, HKONT and KOSTL stay text because every value is a zero-padded code that reading as a number would destroy.

Point it at a directory

Pass --out-dir with one or more files or directories and every .qvd becomes a .parquet of the same name. Add --recursive to descend into subdirectories, and --keep-tree to write each output in the same subfolder its input has, so a folder of table folders converts into a folder of table folders: qvd-delta\VBAK\VBAK.qvd becomes parquet-delta\VBAK\VBAK.parquet, ready to be promoted as one dataset per table.

qvd2parquet --out-dir ./parquet --log run.jsonl ./qvds
qvd2parquet: converting 4 file(s)
qvd2parquet: ok   qvds/products.qvd -> parquet/products.parquet (77 rows, 9 columns, 4.6 KiB)
qvd2parquet: ok   qvds/sales.qvd -> parquet/sales.parquet (1,000 rows, 7 columns, 7.7 KiB)
qvd2parquet: FAIL qvds/truncated.qvd: no XML header terminator (0x00) found: not a QVD file?
qvd2parquet: ok   qvds/stock.qvd -> parquet/stock.parquet (120 rows, 19 columns, 7.3 KiB)
converted 3/4 file(s) in 34ms: 1,197 rows, 19.6 KiB

A bad file in the middle does not stop the run. Every input is attempted, failures are listed at the end, and the exit code reports the most actionable one: a schema policy error you can fix outranks a generic read error.

Folder conversion is built in rather than left to a shell loop for one reason. --file-workers converts several files at once and divides the decode workers between them, so the total stays inside the automatic worker budget. Four separate processes would each start their own full set of workers and oversubscribe the machine fourfold.

A wildcard in the last element of a path is expanded by the tool itself, so ./qvds/CE*.qvd works the same in cmd.exe and PowerShell as in bash. --include-files and --exclude-files narrow what a directory contributes, which is what --recursive needs, and a pattern that reached nothing is reported rather than passing silently.

Only what changed

Over a folder re-extracted nightly, most inputs have not changed since the last run. --skip-up-to-date leaves those alone:

qvd2parquet: skip qvds/A057.qvd (up to date)
qvd2parquet: skip qvds/BSEG.qvd (up to date)
qvd2parquet: stale qvds/CE10500.qvd (input modified since, 2026-09-15T02:10:44.51Z, was 2026-09-14T02:09:58.07Z)
qvd2parquet: ok   qvds/CE10500.qvd -> parquet/CE10500.parquet (20,589,661 rows, 213 columns, 1.8 GiB)
converted 1/3 file(s) in 14m22s: 20,589,661 rows, 1.8 GiB; 2 skipped

It is not a timestamp comparison. Whether the .parquet is newer than the .qvd answers the wrong question: it cannot see that the conversion options changed, a re-extracted file copied with its timestamps preserved looks older than the output it should replace, and two clocks on a network share need not agree. The run keeps a manifest under --out-dir instead, and a file is skipped only when this exact run already produced it: the same input at the same size and timestamp, the same options fingerprint, and the same output still in place. A batch that is stopped halfway resumes on the next run, and a lost manifest can at worst repeat work. A file that is not skipped says which check failed before it starts, so a folder that converts itself every night despite the flag can be diagnosed from the run's own output.

Stable types across daily deltas

A type inferred from the values is exact for the file it came from, and that is the catch with daily delta extracts. A day of small amounts writes decimal(5,2), a day of large ones decimal(9,2), a day of whole numbers int64, and a query engine that promotes the folder as one dataset finds files that disagree. In Dremio the folder shows up empty.

So a folder can carry its own schema. A qvd2parquet-schema.json next to the QVDs pins the types of every file in that folder, and one run over a tree of table folders gives each table its own pins with no extra flag. Generate it from the main table's Parquet and the deltas get exactly the main table's types, so merging them into it never has to cast. It lives beside the inputs rather than the outputs, because the output folder is the one the engine reads as data.

qvd2parquet: schema: pins from qvd-delta\BSEG\qvd2parquet-schema.json
qvd2parquet: schema: DMBTR: pinned to decimal(18, 2) by qvd-delta\BSEG\qvd2parquet-schema.json
qvd2parquet: WARNING: 1 column(s) not pinned by qvd-delta\BSEG\qvd2parquet-schema.json, so their
  types are inferred from this file and may differ from day to day: ZZNEW1

A pin is what the column is. A value with more decimals than its pinned scale is rounded to it the way Qlik displays it and counted in the log, so one odd amount cannot stop a folder of deltas. A value the pin cannot hold at all, such as cents in a column pinned to int64 or more digits than its precision, fails that file with a schema policy error instead of writing one that disagrees with the rest. A field SAP adds to the table is named in a warning on its first day, before it has had a chance to drift. Each folder's file is read once per run, and --skip-up-to-date reconverts only a folder whose schema was added, edited or removed.

Exact decimals

SAP figures are not roughly right, so nothing here is carried through a double. MONEY and FIX are always written as Parquet decimals, with values held as scaled integers end to end. Scale comes from the QVD's NumberFormat/nDec, or is inferred from the display strings when it is absent.

QlikView often declares a price as a plain REAL, where the header carries no usable scale at all. By default those columns are promoted to an exact decimal whose scale is derived from the values themselves: the smallest scale at which every value is exactly representable, up to nine decimals. Pure-integer columns stay int64, since decimal(p,0) would gain nothing.

qvd2parquet: schema: Einkaufspreis: REAL with 75 double symbols promoted to
  decimal(5,2); scale 2 inferred from values

A value that does not fit its declared scale is rounded the way Qlik's number format displays it: the digits the value reads as are rounded half up, whatever the double behind them holds. 1.005 is stored as 1.00499999… and written as 1.01, and -2.345 as -2.34, as Qlik shows them, where rounding the double would give 1.00 and -2.35. The rounding is counted and reported rather than done quietly. --decimal-strict turns it into a failure naming the column and the offending value, for a pipeline where an unexpected precision change has to stop the job; a pinned scale always rounds. A value is never dropped: turning an inexact number into a null would lose data no later check could recover.

Types are resolved, not guessed

The Parquet schema is resolved only after every selected column's symbol table has been read and profiled. That makes mixed-type behaviour explicit, instead of producing an unstable schema that depends on which rows happened to be seen first.

QVD fieldParquet type
INTEGER with integer symbolsint64
REAL with double symbolsdecimal128(p, s), scale inferred
MONEY, FIXdecimal128(p, s), never float64
DATEdate32
TIMESTAMPtimestamp[us]
TIMEtime32[ms]
ASCII or text-only symbolsutf8

One schema: line is printed per output column, explaining exactly why each type was chosen, and --schema-report writes the same reasoning as JSON. That is the first thing to read when a column does not resolve the way you expected.

Qlik duals

A Qlik dual pairs a number with a display string, and often that string is only the number formatted: 1.234,56 beside 1234.56. Writing it would duplicate the numeric column, so the default drops it. When the string carries something the number does not, such as Open beside 1, it is kept as a separate text column and the reason is stated:

schema: Status: INTEGER with 3 integer symbols, written as int64; 3 of 3 display
  strings carry text the number does not (e.g. "Open" beside 1), so they are kept
  in "Status__text"

A single odd value is enough to keep the column, because the default errs towards preserving data. --dual overrides the call in either direction, and --mixed decides what happens to a column mixing numbers with unrelated text: fail with the counts of each symbol kind, write the whole column as text, keep numerics numeric, or split the two sides into separate columns.

One shape resolves on its own, because nothing is actually being decided. When the numeric symbols are integers and every symbol carries its own display string, the file already states the text for every value, so the column is written as text without inventing a rendering for anything. A part number holding 0901 beside the number 901 is a code, not a quantity, and reading it as 901 would not survive a round trip.

Timestamps stay what Qlik wrote

Qlik stores dates and times as serial day numbers. A serial names no timezone. It is a bare wall-clock reading, and which zone it was recorded in is simply not in the file.

So the default converts nothing. --timezone=none writes the wall clock as-is with no timezone on the column, asserts nothing the QVD does not say, and produces a byte-identical file whatever machine runs the conversion. No zone is guessed behind your back.

When you do know the provenance, naming it is the mode that earns its keep for Parquet. An IANA name asserts that the wall clocks were recorded in that zone and converts them to true instants, which is what makes ordering across a DST change, joins against other instant data, and rendering in a consumer's own zone come out right.

--timezone=none              timestamp[us]            2016-03-01 00:00:00
--timezone=UTC               timestamp[us, tz=UTC]    2016-03-01 00:00:00Z
--timezone=America/Chicago   timestamp[us, tz=UTC]    2016-03-01 06:00:00Z
--timezone=Asia/Tokyo        timestamp[us, tz=UTC]    2016-02-29 15:00:00Z

Why the default is not Local: choosing a zone is a claim about provenance that only the person running the conversion can make, and Local makes it accidentally, using whatever zone the converting machine sits in. Real data shows why that matters. In the Chicago taxi QVDs, 13 March 2016 runs 01:45, then 03:00, with the 02:00 hour absent because the local clock skipped it, which also settles what those readings are. Plenty of local-time data contains that missing hour instead: generated master calendars enumerate every slot whether the zone had it or not, and extracts from systems that do not observe DST write wall clocks the target zone never had. A zoned conversion silently relocates every one of those.

Know your target engine. Dremio renders the stored value verbatim and applies no session zone in either direction, so both modes are stable there. DuckDB re-renders an instant in whatever zone the session is set to. Parquet cannot record a timezone name at all, which is why every zoned mode writes tz=UTC: once the readings are on the timeline, what is stored is a UTC instant, and naming UTC keeps every reader agreeing.

A quality gate, not a hope

Every conversion is checked, because a conversion nobody checked is not one anybody can trust. --quality-gate validates the written Parquet against metrics collected from the values the converter actually produced, and it always reads the temporary file before the final rename, so a failed gate never leaves a final-looking output behind.

ModeWhat it checks
basicthe file opens; row count, column names and types match the resolved schema; per-column null counts match
numericeverything in basic, plus sum, min and max per numeric, decimal, date, timestamp and time column
full (default)everything in numeric, plus order-independent SHA-256 value fingerprints per column
rereadeverything in full, plus a second, independent read of the QVD compared against the first, byte for byte

Integer, decimal and date aggregates are compared exactly, with decimal sums using scaled-integer arithmetic and no floating-point tolerance. The full fingerprint is a multiset digest rather than an ordered stream hash, and nulls are marked explicitly so a null never collides with a zero or an empty string.

What none of those modes can question is the read itself. A record byte read wrong yields a different symbol index, that index yields a value that is entirely well formed, and the Parquet then faithfully contains it. reread reads every symbol table and record chunk twice as the conversion goes and stops at the first range that does not match, naming the offset and the two bytes it saw. Nothing read from such a file is trusted and no output is kept.

The default is not free: the whole output is read back, and full digests every value as it is written. On a 213-column, one-million-row fixture that is 27 seconds against 10 without a gate, and the read-back splits across workers, from 61 seconds single-threaded to 9 at eight. Name numeric or basic explicitly when throughput matters more than the fingerprints.

A run you can query and audit

--log writes JSON Lines: one record per file, then a summary. The format is chosen so a finished run can be queried rather than read.

duckdb -c "select status, count(*), sum(rows) from read_json_auto('run.jsonl')
           where type='file' group by 1"
duckdb -c "select input, error from read_json_auto('run.jsonl') where status='failed'"

Each record carries row and column counts, output size, elapsed time, throughput, and the quality gate's verdict with any errors, so a batch can be audited without opening every per-file report. A record is written the moment its file finishes, so a batch that is stopped or killed leaves the lines for what it did convert, and a log without a summary line is a run that did not finish. --console-log copies the screen output to a file as it is printed, for the scheduler that keeps no terminal. Exit codes name the failure class, which is what makes the tool usable from a scheduler:

CodeMeaning
0success
1CLI usage error
2unsupported QVD feature
3schema or type policy error
4input read or decode error
5output or write error
6quality gate failure
7cancelled by Ctrl-C or SIGTERM

Ctrl-C stops the conversion at the next chunk boundary, drains what is in flight, removes the temporary output and exits 7. An existing file at the output path is untouched, since the rename happens last, and a cancelled run is reported as cancelled rather than as bad data. A second signal terminates immediately.

The Parquet file is the only thing written to stdout's usual place: the identification banner, the per-column schema decisions, progress and the final summary all go to stderr, and stdout stays empty, so the tool composes safely in pipelines and shell substitutions.

Inspect before you convert

--inspect reads the XML header and symbol tables, prints the schema a conversion would produce, and exits without touching the record area, so its cost is independent of row count. On a 29 MiB, 5-million-row file it reads 7.9 KiB and finishes in 0.01s, against 1.58s for the full conversion. The one exception is --encoding auto, described below, which makes inspect read the sampled record windows it measures; without it, inspect touches nothing but the header and the symbol tables. Every type policy flag applies, so what you see is what a conversion would write, which makes it a cheap pre-flight check in a pipeline. Each column is shown with its value range, rendered in its resolved type, and a decimal whose widest value already fills most of its precision is flagged before a later extract overflows it.

SAP field names

QVD field names from SAP extracts are often composite, packing the table, the technical name and a description into one string. --exclude strips QlikView's internal key fields by wildcard, and --field-regex splits the rest:

qvd2parquet \
  --exclude '%*' \
  --field-regex '^[^-]*-\|\|-(?P<name>[^-]*)-\|\|-(?P<comment>.*)$' \
  A057.qvd a057.parquet

The result carries the description as Parquet field metadata and keeps the original QVD name, so nothing is lost:

DATBI  int64             {"comment": "Ende Gültigkeit", "qvd.field": "A057-||-DATBI-||-Ende Gültigkeit"}
KBETR  decimal128(4, 2)  {"comment": "Betrag",           "qvd.field": "A057-||-KBETR-||-Betrag"}

A rule that keeps only the technical part of a name is exactly the kind that collapses two fields onto one. By default that is a schema policy error naming both source fields, because it is usually a rule coarser than intended. --duplicate-names=suffix keeps both instead, writing the second as DATBI_2 with its own comment and provenance intact.

A catalog for the engine that has no column comments

A field comment reaches the Parquet file as Arrow field metadata. pyarrow, polars and Arrow-Go decode it. Query engines mostly do not, and Dremio has no column description field at all: DESCRIBE TABLE returns nine columns and none of them is a comment. So the comment reaches a reader there only as data, and --catalog-out writes that data: one Parquet row per output column, for the whole run.

qvd2parquet --out-dir out --catalog-out catalog.parquet \
  --field-regex '^[^-]*-\|\|-(?P<name>[^-]*)-\|\|-(?P<comment>.*)$' \
  extracts/
source  source_table  ordinal  column_name  comment          parquet_type
qvd     A057          1        DATBI        Ende Gültigkeit  int64
qvd     A057          2        KSCHL        Konditionsart    utf8
qvd     A057          3        KBETR        Betrag           decimal(4, 2)

Every row also carries the run, the tool version, the source file and its row count, the Qlik type, nullability, the symbol count, the value range and the reason the type was chosen. The schema is fixed and every field is written on every row, so a query does not break on the one run where nothing was commented. The catalog is a table: promote it as a dataset and join it against INFORMATION_SCHEMA.COLUMNS, and it answers which columns are undocumented and which exist in the engine but not in the last conversion, which is schema drift arriving as a row rather than as a surprise.

A catalog is merged into, not replaced. A nightly job converts the tables that changed, so a catalog holding only the last run would describe one table on Tuesday and a different one on Wednesday. Rows are keyed by table and column: a column the run describes is updated in place, a new one is added, a table the run did not touch keeps every row it had, and a column a table no longer has is kept with the run that last saw it. Nothing is ever removed, which is what makes drift queryable. A run that forgot the flag is not lost either: --catalog-scan builds the same table from the footers of Parquet files already written, one seek per file however many rows they hold.

Encodings for SAP keys, measured

Every column is written with a dictionary by default, which is right for the columns a QVD is usually full of. It is worth nothing on a Qlik composite primary key, where every row has its own value: the dictionary overflows and the column ends up as raw bytes with only the compressor working on it. --encoding pins such a column to something better, with wildcard rules that cover a folder of SAP tables whose keys are named per table:

qvd2parquet --encoding '%*_PKEY=delta_byte_array' CE10500.qvd ce10500.parquet

Whether delta_byte_array pays depends on the order the rows arrive in, since it stores each value against the one before it. On a key in document order it cuts the column to about a third; on the same values shuffled it saves nothing, and nothing in the symbol table reveals which case a file is. So --encoding auto measures it: sampled rows are written through the real writer twice, once as the run would today and once with each candidate, and the compressed size of the column chunk decides.

This is the one case where --inspect reads records: three windows of 100,000 consecutive rows at the head, the middle and the tail, or the whole file when it has 300,000 rows or fewer. The header-only preflight above stays header-only without this flag.

qvd2parquet --inspect --encoding auto CE10500.qvd
Columns that would compress better with a different encoding:
  %CE10500_PKEY  delta_byte_array, measured 31% of current size on 300,000 sampled rows
A conversion with --encoding auto writes them that way; pin them instead with --encoding "%CE10500_PKEY=delta_byte_array".

That run wrote 1.8 MiB where the default wrote 6.2 MiB, which is what the sample predicted before the conversion started. A candidate has to measure at 80% of the current size or better to be adopted, on a shuffled key nothing is adopted and the run says so, and an explicit rule always wins over a measurement.

Performance

The record area is fixed-width, so once the symbol tables are read it can be split into contiguous row ranges and decoded concurrently. Each worker owns its Arrow builders, reads its byte range directly, and emits one Arrow record plus chunk-local quality metrics, while a single writer goroutine feeds the Parquet writer.

Measured on an Apple M3 Max with 16 cores over a 200k-row fixture with integer, high-cardinality string, decimal, date and nullable double columns: decoding reaches 18.4M rows/s at the default worker count and 28.6M with one worker per CPU, and the full pipeline including zstd-compressed Parquet writing runs at 3.3M rows/s. zstd is both the smallest and the fastest option here, which is why it is the default.

Nothing in the pipeline scales with total row count. The default batch size is chosen by cells rather than rows, about two million per batch, so in-flight memory stays put instead of growing with width, and the default worker count is one per two CPUs, because on a wide file in-flight memory is the binding constraint, not decode. On a 213-column, 20.6-million-row SAP extract on a 16-core Xeon, the conversion held 23.9k rows/s end to end with the full quality gate on, at 9.3 GB resident, where the 1.0 release had climbed to 36 GB and slowed as it did.

Row order is preserved. Decoding runs in parallel, but records reach the writer in chunk order, so the Parquet file holds the QVD's rows in their original order whatever the worker count, and the output is byte-comparable across worker counts. That is also what makes row-group statistics useful: a row group's min and max only bound the rows it holds when those rows are contiguous in the source, which is what lets an engine skip row groups over a sorted key.

Install

Download the archive for your platform, unpack it, and put qvd2parquet on your PATH. The binaries are pure Go and statically linked, so they have no runtime dependencies: Linux, Windows and macOS, on both amd64 and arm64. Verify the download against the published SHA256SUMS.

shasum -a 256 -c SHA256SUMS --ignore-missing

Or build from source with Go 1.25 or newer:

go install github.com/ralforion/qvd2parquet/cmd/qvd2parquet@latest

From version 2.0.0 the CLI surface and the conversion defaults are stable: a flag will not be removed or change its meaning, and a default will not change what an existing file converts to, outside a major bump. New behaviour arrives behind a new flag or a new value for an existing one. 2.0.0 was the major bump where the defaults were re-chosen for wide files. 3.0.0 is the next one: no flag changed, but decimals are now rounded the way Qlik displays them, which changes values already written, so the first --skip-up-to-date run after upgrading reconverts every file once. The promise restarts from it. The tool is released under the Apache License 2.0.

Built for a Qlik to Dremio migration

A QVD layer is often a decade of accumulated business logic, and moving it to a lakehouse is where that history either survives or quietly degrades. Stringified numbers, doubles standing in for currency, zero-padded account codes read as integers, and timestamps shifted by whichever machine ran the job are the failures that show up months later in a reconciliation, not on the day of the migration.

That is the job this tool was written for: land a Qlik estate as Parquet in object storage, ready for Dremio to query as Iceberg tables or promoted datasets. The defaults follow from that target. Exact decimals because DMBTR has to reconcile against SAP. Text preserved for zero-padded codes because they are keys, not quantities. Timestamps left as the wall clock Qlik recorded, because Dremio renders the stored value verbatim and applies no session zone, so an unnecessary conversion would be visible in the data and invisible in the schema.

Getting the physical data out correctly is the first half of that move. The second is the business meaning that lived in the Qlik load scripts, which is what the OrionBelt Semantic Layer is for: define dimensions, measures and metrics once in version-controlled YAML, and compile them into correct SQL against Dremio, or against whatever else the lakehouse turns out to hold.

Frequently Asked Questions

Does this work for a Qlik to Dremio migration?

That is what it was built for. It lands a QVD estate as Parquet in object storage for Dremio to query as Iceberg tables or promoted datasets. Dremio renders a stored timestamp verbatim and applies no session zone in either direction, so both timezone modes are stable there, which is also why the default asserts nothing the QVD does not say.

How do I convert a Qlik QVD file to Parquet?

Download the binary for your platform and run qvd2parquet input.qvd output.parquet. There is no runtime to install. To convert a whole folder, pass --out-dir with one or more files or directories, and add --recursive to descend into subdirectories.

Are Qlik MONEY and decimal fields converted exactly?

Yes. MONEY and FIX are always written as Parquet decimals and carried as scaled integers end to end, so no step rounds through a double. A plain REAL column holding fractional values is also promoted to an exact decimal by default, with the scale derived from the values. A value that does not fit its scale is rounded the way Qlik displays it and counted, and --decimal-strict makes it fail instead, except in a pinned column, which always rounds.

What happens to timestamps and timezones?

A Qlik serial names no timezone, so by default the wall clock is written as-is with no timezone on the column, and the output is byte-identical whatever machine converts it. Pass --timezone with an IANA name to assert the zone the readings were recorded in and convert them to true instants.

Can it convert a whole folder of QVD files?

Yes. --out-dir converts every QVD in the tree, running several files at once with --file-workers and dividing decode workers between them so the machine stays near one worker per CPU. A failing file is logged, skipped, listed at the end, and reflected in the exit code.

How do I know the converted file is correct?

Every conversion is checked by default. The quality gate reads the written Parquet back before the final rename and compares it against metrics collected during conversion: row counts, types and null counts, exact numeric aggregates, and order-independent SHA-256 value fingerprints. --quality-gate reread also reads the QVD a second time and compares the two reads byte for byte, which is the only check that can catch a source read that went wrong. Name a cheaper mode when throughput matters more.

Does it read encrypted QVD files?

No. It reads standard unencrypted QVD files. It also does not write QVD files or produce nested Parquet output. Row order is preserved: decoding runs in parallel, but records reach the writer in their original order.

How do I get column descriptions into Dremio?

Dremio has no column comment field, so a description stored as Parquet field metadata never reaches a query. --catalog-out writes it as data: a Parquet table with one row per output column, carrying the comment, the QVD and Parquet types, the value range and the run that wrote it. The catalog is merged across runs rather than replaced, and --catalog-scan builds the same table from files converted earlier.

Does a nightly job reconvert the whole folder every time?

Not with --skip-up-to-date. A manifest under the output directory records which input, at which size and timestamp, with which options, produced which output, and a file is skipped only when this exact run already produced it. A re-extracted source, a changed option or a replaced output reconverts it, and the run says which of those it was before the file starts.

Why does Dremio show a folder of daily delta files as empty?

Usually because the files disagree on a column's type. A type inferred from the values fits the file it came from, so one small delta writes decimal(5,2), a larger one decimal(9,2) and a day of whole numbers int64. A qvd2parquet-schema.json next to the QVDs pins the types for every file in that folder; generated from the main table's Parquet, it gives the deltas the main table's schema. A column the file leaves unpinned is reported with a WARNING line, and a value with more decimals than its pin is rounded to it and counted, so every file keeps the pinned type.

Who built qvd2parquet?

Ralfo Becher and RALFORION d.o.o., the team behind the source-available OrionBelt Semantic Layer. It is released under the Apache License 2.0.

Get the binary

qvd2parquet is free and open source, built by Ralfo Becher at RALFORION, the team behind the OrionBelt Semantic Layer. Download it for Linux, Windows or macOS, or get in touch about the rest of the migration.

Download GitHub Contact RALFORION