Skip to content

Retry, resume and local history

Set store= when creating a backend to continue an experiment from another LocalBackend instance or Python process. The SQLite file holds task definitions; result files live in the adjacent <store>.artifacts directory and are verified by checksum when read.

Save a result and read it again

This example completes a run, then restores the same job through a new backend instance. It uses a temporary directory that is removed on exit. Choose a lasting path when saving your own experiments.

"""Restore a saved result through a new backend instance."""

from __future__ import annotations

import json
from pathlib import Path
from tempfile import TemporaryDirectory

from cascaqit import Circuit, LocalBackend, RetryPolicy


def main() -> None:
    """Restore a saved result through a new backend instance."""
    with TemporaryDirectory(prefix="cascaqit-resume-") as directory:
        store = Path(directory) / "runs.sqlite"
        backend = LocalBackend(store=store, seed=7)
        job = backend.run(
            Circuit(1).h(0),
            shots=100,
            retry=RetryPolicy(max_attempts=3),
            idempotency_key="experiment-001",
        )
        result = job.result()
        restored = LocalBackend(store=store).resume(job.job_id)
        same_result = restored.result().stable_hash() == result.stable_hash()
        page = backend.history(limit=20, statuses=("completed", "failed"))
        if not same_result or len(page.entries) != 1:
            raise RuntimeError("A restored result must match the saved result.")
        print(
            json.dumps(
                {
                    "restored_result_matches": same_result,
                    "counts_total": sum(result.counts.values()),
                    "history_entries": len(page.entries),
                    "restored_status": restored.status().state,
                },
                sort_keys=True,
            )
        )


if __name__ == "__main__":
    main()

Download the full script

python examples/learning/guides/resume_job_en.py
{
  "counts_total": 100,
  "history_entries": 1,
  "restored_result_matches": true,
  "restored_status": "completed"
}

Expect restored_result_matches to be true, one history entry and a total of 100 counts. Restoring a completed job reuses its saved result. This example checks recovery; it does not inject a failure to demonstrate retries.

The complete offline example also saves and restores a Hybrid parameter scan.

Which errors are retried?

An error is retried only if it is a CASCAQit exception marked as retryable or its code is explicitly allowed by RetryPolicy. Validation, compilation, parameter and storage-corruption errors are not retried by default, nor are unclassified Python exceptions. Increasing max_attempts cannot fix invalid parameters.

Without store=, jobs remain in the current process and compute when result() is called. Both resume() and history() require a store.

Settings used on recovery

Recovery uses the program, bindings, seed, shots, simulation options, noise, observables, retry policy and scan order saved at submission. Defaults on the new backend do not replace them.

An explicit idempotency_key prevents duplicate submissions of the same definition. Reusing a key with different execution settings raises an error. Runs submitted without a key remain independent.

Completed scan children are reused after checksum verification. Pending or interrupted children continue under the saved retry policy, and the aggregate returns items in scan-index order. See the parameter scan guide for scan setup.

Query and remove history

history() filters by status, program kind, job kind and UTC creation time. A pagination cursor only works with its original store and filters. Each entry contains identifiers, state, attempt or item counts, timestamps and references to results or errors. Listing history does not load the full result files.

Jobs do not expire automatically, and there is currently no bulk cleanup API. When the results are no longer needed, remove the SQLite file and its matching <store>.artifacts directory together.

SQLite coordinates execution attempts on one host. It provides neither distributed scheduling nor an exactly-once guarantee across hosts. Attempt timeouts require the running code to cooperate; they cannot forcibly stop NumPy, SciPy or native extensions. Local persistence does not connect to hardware or cloud services.

For the API details, see LocalBackend and RetryPolicy.

中文版

SDK 1.0.8a · `8b227bff`