Proto commits in lakesoul-io/LakeSoul

These 10 commits are when the Protocol Buffers files have changed:

Commit:8465437
Author:Xu Chen
Committer:GitHub

[Metadata] snapshot, tag and pin-aware retention (#925) * [Docs] embodied: snapshot/tag/retention design (P0) Records the reviewed design: table snapshots, tags as retention pins, pin-aware Flink clean job, blob reference model (R1: independent packs plus a per-data-file .blobref sidecar and orphan GC), and manifest sibling tables. Branching is out of scope. Also captures the explicit-timezone rule for timestamp time travel. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Metadata] add snapshot, tag and pin tables (P0 phase A) table_snapshot/snapshot_commit/table_snapshot_tag plus pinned flags on data_commit_info and partition_info. Idempotent migration V4000001__snapshot_tag.sql; meta_init.sql updated for fresh installs. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Python] time travel by timestamp with explicit time zones LakeSoulScan.options(timestamp=..., time_zone=...) resolves the latest version of every partition at or before a point in time via the existing ListPartitionByTableIdAndTimestamp DAO (metadata_client get_all_partition_info_as_of). date/datetime/ISO-8601 values with a time zone are used as-is; values without one require an explicit time_zone and raise otherwise, so a naive value can never be interpreted silently. Timestamps are PG-server commit times in UTC epoch ms. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Python+Rust] table snapshots and tags (P0 phase A) New metadata DAOs for table_snapshot/snapshot_commit/table_snapshot_tag. Creating a snapshot freezes every partition's version and commit ids in one statement and pins the referenced commits/versions; dropping a snapshot unpins what is no longer referenced and refuses while a tag still points at it. Python gains create_snapshot/create_tag/list_snapshots/list_tags/drop_tag/ drop_snapshot and scan options snapshot=/tag= (mutually exclusive with timestamp=), resolving files straight from the frozen commit ids instead of relying on old partition_info rows. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Flink] clean job skips pinned commits (P0 phase B, part 1) The compaction cleanup path now adds pinned = false to the file-selection query, the data_commit_info delete and the partition_info delete, so commits/versions frozen in a snapshot or tag survive the Flink clean job. The TTL partition-drop path, discard-file handling and the purge API still need the pin state exposed through the Java DAOs; local Maven compilation failed with an environment OOM, so CI must verify this change. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Metadata] expose pinned on PartitionInfo and DataCommitInfo Adds the pin flag to both protos and to every partition/data-commit select (including the aliased join queries), so Java DAOs can check pin state when deciding whether to drop a partition or delete discard files. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Flink] clean job honours pinned metadata (P0 phase B, part 2) TTL partition drop now skips partitions that have any pinned version, and the timestamped cleanup filters pinned versions out; the partition_info delete keeps pinned rows on both the Rust and JDBC paths. Discard file deletion resolves pinned paths through data_commit_info.pinned and skips them. Lakesoul-common and lakesoul-flink compile locally with JDK 11 (maven build cache disabled; the cache extension OOMs on this repo). Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Python] purge API with grace period (P0 phase B, part 3) Adds exec_update plumbing (MetaDataClient::execute_update_raw, PyO3 exec_update, NativeMetadataClient.exec_update) and catalog.purge(): it removes partition versions that are older than the grace period, not pinned by any snapshot/tag and not the latest version, deletes commits/files only when they are no longer referenced by any live version, and deletes blob packs together with their data file. DeleteDataCommitInfoByTableIdAndPartitionDescAndCommitIdList now filters pinned = false like the other cleanup statements. Tests cover pinned/latest retention, removal after unpinning, and the cleanup SQL statements directly against pinned rows. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Spark] skip pinned files when recording discard entries (P0 phase B, part 4) Compaction no longer writes files that a snapshot or tag still pins into discard_compressed_file_info, so the Flink clean job never even considers them; the delete-side guard remains as defence in depth. Verified by compiling lakesoul-common and lakesoul-spark locally (JDK 11, maven build cache disabled). Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Common] PG tests for pinned cleanup guards PinnedCleanupTest inserts pinned/unpinned partition versions and commits directly in PostgreSQL and checks getPinnedFilePaths and deleteMetaPartitionInfo: pinned partitions are never dropped, unpinned ones are. Fixes found by the test: the JDBC commit-list query now casts placeholders to uuid, and both JDBC row mappers populate the pinned column. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Flink] clean job mini-cluster test for pinned discard files NewCleanJob exposes buildPipeline(env, parameter) so tests can run the real pipeline on a mini cluster. CleanJobPinnedTest starts a logical-replication PostgreSQL in docker (schema from script/meta_init.sql), inserts a pinned data commit plus discard entries for a pinned and an unpinned file, runs the pipeline with a 2s expiry and asserts the pinned file and its discard row survive while the unpinned ones are removed. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] compaction clean e2e workflow and env-provided PG test CleanJobPinnedTest no longer starts PostgreSQL itself: it expects a logical-replication instance with the LakeSoul schema (CI service or local docker), configured through LAKESOUL_PG_* and verified with show wal_level. The new compaction-clean-e2e workflow provides postgres:14 with wal_level=logical plus RustFS, builds the native libraries, initializes the schema and runs the test; the tag/blob and multi-round compaction scenarios will extend the same workflow. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] compaction clean e2e as a runnable script script/ci/compaction_clean_e2e.sh deploys the docker compose environment (logical PostgreSQL, RustFS, Flink cluster), submits the clean job to the Flink cluster and starts the Spark NewCompactionTask in the background, then runs script/ci/compaction_clean_e2e.py which writes several data rounds and checks the normal, tagged and blob scenarios. The GitHub workflow only builds the artifacts and calls the script, so the same flow runs locally. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] fix compaction clean e2e after local run Verified script/ci/compaction_clean_e2e.sh --down locally. Fixes: - wait for PostgreSQL before enabling logical replication - create the RustFS bucket with boto3, the bundled rc client fails on 1.0.0-beta.3 - clear the proxy environment injected by the docker client so S3A reaches RustFS - use the compaction task argument syntax understood by ParametersTool and enable the new compaction path - run a single named compaction container and remove it on exit - restore script/meta_init.sql through an exit trap - clean the expired data immediately (dataExpiredTime 0) and assert file deletion from the object store instead of cached partition metadata Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] fix disabled crate build and formatting lakesoul-datafusion initializes DataCommitInfo without the new pinned field, which broke clippy, rust_ci, run-pytest and hash consistency jobs; add the field. Also apply rustfmt, ruff and google-java-format to the changed files so the format check passes. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] match google-java-format continuation indent Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [Metadata] add pinned to partition partial-filter queries ListPartitionByTableIdAndFilterCondition selected eight columns while the row mapper reads pinned from index 8, breaking merge with partial partition filters; the Java fallback queries had the same gap. Also pull the Spark image before starting the compaction container and allow a longer cleanup wait in the compaction clean e2e script, the CI runner pulled the image while the assertions were already polling. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] provide logical replication for the pinned clean test CleanJobPinnedTest requires wal_level=logical; add the postgres service option for the flink test job and skip the test when the environment does not provide logical replication instead of failing. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] reuse the rust cache and build natives once The compaction clean e2e job rebuilt the rust crates for the native libraries and again through maturin. Follow python-ci: one cargo build producing the io/metadata/python libraries, copy liblakesoul_python.so into the python package and install dependencies with --no-install-project. Use the native-glibc217 shared rust cache so PR runs restore the main cache instead of compiling the full dependency tree. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] match the zigbuild cache in compaction clean e2e native-glibc217 is the cargo-zigbuild cache populated by native-build and maven-test; a plain cargo build would not hit it. Use the same cargo zigbuild --target x86_64-unknown-linux-gnu.2.17 --all-features command and install Zig plus cargo-zigbuild. Build the python extension separately without the ci-unify feature so lakesoul-datafusion is not compiled for this job. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] follow python-ci for the e2e native build Build every package with a single cargo build as python-ci does, and use the same rust-cache setup without a shared key or cargo-zigbuild. The python extension is copied into the package from the plain release output, and uv only installs dependencies. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] use zigbuild and the native-glibc217 cache for the e2e The Flink and Spark containers need glibc 2.17 native libraries, so keep cargo-zigbuild for lakesoul-io-c/lakesoul-metadata-c with the native-glibc217 shared rust cache, and build the python extension separately with a plain cargo build that does not need ci-unify. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] build every native artifact with one zigbuild Build lakesoul-io-c, lakesoul-metadata-c and lakesoul-python in a single cargo zigbuild invocation and copy all three libraries from the same target directory, so python no longer triggers a second compile. Use the hdfs feature of lakesoul-io-c to match the maven-test feature set, pin Zig 0.14.1 and set save-if false because maven-test is the only writer of the native-glibc217 rust cache. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] load libhdfs for the hdfs native libraries The native-glibc217 build enables the hdfs feature, so the shared libraries need libhdfs at load time: run the driver with the Hadoop native directory on LD_LIBRARY_PATH and do the same for the Flink and Spark containers. Revert the postgres service option, GitHub parses it as docker --cpu-shares; the pinned cleaning test skips when the service is not configured for logical replication. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> * [CI] add libjvm to the loader path for the e2e driver The hdfs feature links both libhdfs and libjvm, and the python driver runs outside a JVM, so JAVA_HOME/lib/server must be on LD_LIBRARY_PATH next to the Hadoop native directory. setup-java already provides the JDK; the containers do not need it because their JVM has libjvm loaded. Co-Authored-By: opencode <noreply@opencode.ai> Signed-off-by: chenxu <chenxu@dmetasoul.com> --------- Signed-off-by: chenxu <chenxu@dmetasoul.com> Co-authored-by: chenxu <chenxu@dmetasoul.com> Co-authored-by: opencode <noreply@opencode.ai>

Commit:d058215
Author:mag1c1an1
Committer:GitHub

[Build] Foundation plumbing for 4.0.0: build-info crate, proto rename, unified versioning (#825) * chore: add buildinfo and use new version Co-Authored-By: glm-5.2 <noreply@zhipuai.cn> Signed-off-by: mag1cian <mag1cian@icloud.com> * fix: jar version in python ci Signed-off-by: mag1cian <mag1cian@icloud.com> --------- Signed-off-by: mag1cian <mag1cian@icloud.com> Co-authored-by: glm-5.2 <noreply@zhipuai.cn>

Commit:1a46381
Author:mag1c1an1
Committer:GitHub

[Metadata] Use Arrow IPC metadata schema (#800) * vortex: use projection pushdown Signed-off-by: mag1cian <mag1cian@icloud.com> * metadata: schema use arrow ipc Signed-off-by: mag1cian <mag1cian@icloud.com> * [Rust] Fix metadata branch clippy Implement the new DataFusion MemoryPool requirements for LoggedMemoryPool and remove a redundant as_ref in a test assertion. Co-Authored-By: Codex <noreply@openai.com> Signed-off-by: mag1cian <mag1cian@icloud.com> --------- Signed-off-by: mag1cian <mag1cian@icloud.com> Co-authored-by: Codex <noreply@openai.com>

Commit:10bae98
Author:du
Committer:GitHub

[Spark] Implement new compaction strategy (#585) * new compaction Signed-off-by: fphantam <dongf@dmetasoul.com> * remove debug info and change some params Signed-off-by: fphantam <dongf@dmetasoul.com> * remove debug code Signed-off-by: fphantam <dongf@dmetasoul.com> * new compaction Signed-off-by: fphantam <dongf@dmetasoul.com> * add License Signed-off-by: fphantam <dongf@dmetasoul.com> * not use DynamicBucket if table hashBucketNum not change in new compaction task Signed-off-by: fphantam <dongf@dmetasoul.com> * remove useless code Signed-off-by: fphantam <dongf@dmetasoul.com> * fix new compaction bug Signed-off-by: fphantam <dongf@dmetasoul.com> * add CompressDataFileInfo for midle file info Signed-off-by: fphantam <dongf@dmetasoul.com> --------- Signed-off-by: fphantam <dongf@dmetasoul.com>

The documentation is generated from this commit.

Commit:e15c8e7
Author:Ceng
Committer:GitHub

[Flink] add arrow datastream source (#496) * add lakesoul arrow source Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix flink writer path Signed-off-by: zenghua <huazeng@dmetasoul.com> * support inferring schema Signed-off-by: zenghua <huazeng@dmetasoul.com> * tmp commit Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix regression test Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix reading tail recordbatch Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix stup dao sql Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix dynamic schema Signed-off-by: zenghua <huazeng@dmetasoul.com> * optimize imports Signed-off-by: zenghua <huazeng@dmetasoul.com> * split DataFileInfo with file_cols Signed-off-by: zenghua <huazeng@dmetasoul.com> * optimize import Signed-off-by: zenghua <huazeng@dmetasoul.com> * rebase on main Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix regresion test Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix regresion test Signed-off-by: zenghua <huazeng@dmetasoul.com> --------- Signed-off-by: zenghua <huazeng@dmetasoul.com> Co-authored-by: zenghua <huazeng@dmetasoul.com>

Commit:ba75914
Author:zenghua
Committer:zenghua

well define DataFileOp Signed-off-by: zenghua <huazeng@dmetasoul.com>

Commit:8749d2e
Author:Ceng
Committer:GitHub

merge native-io and native-metadata modules (#318) Signed-off-by: zenghua <huazeng@dmetasoul.com> Co-authored-by: zenghua <huazeng@dmetasoul.com>

Commit:cc285cc
Author:Ceng
Committer:GitHub

[Native-Metadata] Rust implementation of DAO layer (#294) * POC of native-metadata-jni Signed-off-by: zenghua <huazeng@dmetasoul.com> * Native DAO Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix git action for native-metadata Signed-off-by: zenghua <huazeng@dmetasoul.com> * cleanup code Signed-off-by: zenghua <huazeng@dmetasoul.com> * fix after rebase main Signed-off-by: zenghua <huazeng@dmetasoul.com> --------- Signed-off-by: zenghua <huazeng@dmetasoul.com> Co-authored-by: zenghua <huazeng@dmetasoul.com>

Commit:8f3778b
Author:Xu Chen
Committer:GitHub

[Project] add/change spdx header to all files (#278) Signed-off-by: chenxu <chenxu@dmetasoul.com> Co-authored-by: chenxu <chenxu@dmetasoul.com>

Commit:83bd0df
Author:Ceng
Committer:GitHub

[Metadata] refractor metadata entity with protobuf code-gen (#276) * refractor metadata entity with protobuf Signed-off-by: zenghua <huazeng@dmetasoul.com> * add git action setup protoc Signed-off-by: zenghua <huazeng@dmetasoul.com> * bugfix for regression test Signed-off-by: zenghua <huazeng@dmetasoul.com> * Delete lakesoul-common/src/main/java/com/dmetasoul/lakesoul/meta/entity directory * add entity comments && fix compile option Signed-off-by: zenghua <huazeng@dmetasoul.com> --------- Signed-off-by: zenghua <huazeng@dmetasoul.com> Co-authored-by: zenghua <huazeng@dmetasoul.com>