Skip to content

Resume a saved local job without rerunning completed work

Create a persistent Job, discard its original Backend handle, and resume it through a new Backend using the same SQLite store. Complete scan outcomes and use the installed environment.

The circuit prepares H on one qubit and samples eight shots. Its ideal probabilities are balanced, but the lesson's key prediction concerns execution: resuming queued work executes once, and reading the completed result again does not create another attempt.

Restore the frozen definition

"""Persist, resume, and query one local Job with immutable evidence.

SQLite stores a frozen Job definition, attempts, transitions, and a checksum
ResultRef. A new Backend instance can resume the Job on the same host. This is
not a distributed scheduler or cross-host exactly-once guarantee.
"""

from __future__ import annotations

import json
from pathlib import Path
from tempfile import TemporaryDirectory

from cascaqit import Circuit, LocalBackend, ResultIR, RetryPolicy


def main() -> None:
    """Run one durable Job and inspect its restored history."""
    with TemporaryDirectory(prefix="cascaqit-track-history-") as directory:
        store = Path(directory) / "runs.sqlite"
        retry = RetryPolicy(max_attempts=3, backoff_strategy="none")
        backend = LocalBackend(store=store, seed=404)
        job = backend.run(
            Circuit(1, program_id="lesson.platform.durable").h(0),
            shots=8,
            retry=retry,
            idempotency_key="lesson-platform-durable",
        )
        queued_state = job.status().state
        restored = LocalBackend(store=store).resume(job.job_id)
        result = restored.result()
        assert isinstance(result, ResultIR)
        completed_again = LocalBackend(store=store, seed=999).resume(job.job_id)
        saved_result = completed_again.result()
        history = backend.history(limit=10, statuses=("completed",))

        facts = {
            "queued_state": queued_state,
            "completed_state": restored.status().state,
            "attempt_count": restored.status().to_dict()["execution_count"],
            "completed_read_matches": saved_result.stable_hash()
            == result.stable_hash(),
            "attempts_after_completed_read": completed_again.status().to_dict()[
                "execution_count"
            ],
            "counts_total": sum(result.counts.values()),
            "history_entries": len(history.entries),
            "result_ref_present": history.entries[0].result_ref is not None,
            "retry_max_attempts": retry.max_attempts,
        }
    payload = {
        "track": "sdk_platform_engineer",
        "level": "advanced",
        "lesson": "retry_resume_history",
        "facts": facts,
        "boundaries": {
            "hardware_execution": False,
            "cloud_execution": False,
            "network_accessed": False,
            "credentials_loaded": False,
        },
    }
    print(json.dumps(payload, sort_keys=True))


if __name__ == "__main__":
    main()

Download the full script

python3 examples/user/tracks/sdk_platform_engineer/04_advanced_retry_resume_history_en.py

store=... opts into persistence. The first Backend saves the program and execution settings while the job is queued. The second restores them by job_id, then executes on result(). The final Backend uses a different default seed, 999, but reads the already completed result. It does not replace the stored seed or draw another sample.

{
  "boundaries": {
    "cloud_execution": false,
    "credentials_loaded": false,
    "hardware_execution": false,
    "network_accessed": false
  },
  "facts": {
    "attempt_count": 1,
    "attempts_after_completed_read": 1,
    "completed_read_matches": true,
    "completed_state": "completed",
    "counts_total": 8,
    "history_entries": 1,
    "queued_state": "queued",
    "result_ref_present": true,
    "retry_max_attempts": 3
  },
  "lesson": "retry_resume_history",
  "level": "advanced",
  "track": "sdk_platform_engineer"
}

Both attempt counts should be one and completed_read_matches should be true. history_entries=1 counts the saved job, not Backend objects. result_ref_present means the history entry carries a result reference; the history list does not itself contain the full result body.

The example uses TemporaryDirectory, which removes the store and its artifacts when the block exits. It demonstrates recovery between handles while the directory exists. For recovery after this script finishes or in a new process, choose a persistent path and retain both the SQLite file and its matching .artifacts directory. The result sidecars are checked by checksum; a database file alone is not a complete backup.

Retry policy is a limit, not an observed retry

max_attempts=3 permits up to three attempts under the configured policy. This successful example uses one; it does not simulate a transient failure. Recognized retryable errors may cause another attempt. Validation, compilation, parameter errors, store corruption and unclassified Python exceptions are not retried by default. Repeating an invalid input cannot fix it.

The same explicit idempotency_key deduplicates the same frozen definition. Reusing that key with changed execution settings is a conflict, not a request to overwrite the earlier experiment. Use a new key when you intend a new run.

Check recovery and identity

  1. Submit the same circuit and options twice with the same idempotency key. Compare the returned job IDs.
  2. Keep the key but change shots from 8 to 16. Read the error.
  3. Call resume() on a Backend with no store configured. What information is missing?

The first submission pair refers to the same job. The changed shot count raises JOB_IDEMPOTENCY_CONFLICT; the no-store call raises LOCAL_BACKEND_STORE_REQUIRED. A job ID alone does not locate an arbitrary external database.

SQLite coordinates local attempts on one host. This is not a distributed scheduler or a guarantee of cross-host exactly-once execution. Timeouts are cooperative and cannot forcibly interrupt native numerical code. See retry, resume and history, then learn to check capabilities before accepting work.

中文版

SDK 1.0.8a · `8b227bff`