Compare VQE objectives with finite sampling¶
A low energy observed during optimization can be partly a sampling fluctuation. This project compares four objective strategies, independently remeasures candidate parameters, and then checks their energies exactly in a small simulator. You will distinguish estimation error from parameter quality and account for the executions used to obtain both.
Complete VQE and energy estimates, baseline analysis and experimental design. Use the installation environment, then run from the repository root. The default experiment has one qubit, four strategies, three paired repetitions and one SPSA iteration per strategy. It is a local CPU exercise with ideal state evolution; finite shots are its only measurement uncertainty.
Derive a reference you can check¶
Use H = 0.1 I + 0.7 X − 0.4 Z, in an arbitrary energy unit. Since (0.7 X − 0.4 Z)² = 0.65 I, its lowest eigenvalue is E0 = 0.1 − sqrt(0.65) ≈ −0.70623. This ground energy is an independent reference.
The default single-layer circuit applies RY(θ) then RZ(φ) to |0⟩. Its energy is
E(θ, φ) = 0.1 + 0.7 sin(θ) cos(φ) − 0.4 cos(θ).
You can substitute any saved ry_0_0 and rz_0_0 values into this expression. Exact evaluation removes sampling error from the objective. It does not guarantee that one optimizer iteration finds the ground state.
The X and Z terms require two measurement bases. An X measurement applies H before computational-basis measurement; the Z group needs no basis rotation. Each group estimates a ±1 mean from counts. The sampled energy is 0.1 + 0.7 mean(X) − 0.4 mean(Z).
Decide what the budget means¶
Within a repetition, all four strategies start from the same parameters and use the same SPSA seed. Different repetitions use different derived seeds. With perturbation signs Δi ∈ {−1,+1}, SPSA estimates gi = [E(θ+cΔ) − E(θ−cΔ)]/(2cΔi) and updates θi ← θi − a gi. Here a = 0.2, c = 0.15, with one update. Initial, plus, minus and final-center evaluations give four logical objective estimates.
| Strategy | How each logical objective estimate is obtained |
|---|---|
exact |
One state-based energy evaluation |
sampled_single |
One complete two-basis measurement |
sampled_fixed |
Two complete measurements, pooled |
sampled_adaptive |
At least two complete measurements; stop at pooled standard error ≤ 0.08, or after three measurements, subject to the budget |
Each run has a ceiling of 12 objective evaluations, including repeated measurements. It need not consume all 12. The exact and single strategies normally use four; fixed uses eight; adaptive can use eight through twelve. These are equal evaluation ceilings, not equal actual shots or equal total costs.
shots_per_group=128 with allocation="coefficient_l1" sets a total of 2 × 128 = 256 shots per complete sampled energy evaluation. Allocation by coefficient magnitude gives 163 X shots and 93 Z shots. The name of the setting does not imply that both groups receive 128 shots under this allocation rule.
After optimization, the sampled strategies select up to two distinct candidate parameter sets from history and independently measure each twice. In the default runs both candidates are present, so confirmation uses eight Backend executions and 1,024 shots per sampled strategy. Candidate selection cannot be evaluated honestly by reporting only the smallest energy that appeared in the search.
A final computational-basis sample uses another 128 shots per strategy. It describes the selected state in the Z basis; it cannot independently estimate the X term. The benchmark also evaluates each sampled strategy's selected parameters once with an exact simulator. That diagnostic execution is additional cost and would not be available as a free oracle on hardware. Confirmation, final sampling and diagnostics sit outside the 12-evaluation optimization ceiling.
Run and inspect the results¶
python3 examples/learning/projects/finite_shot_vqe.py
Open artifacts/finite-shot-vqe/vqe-en.html. The standard benchmark report shows selected energies, paired comparisons, uncertainty and execution costs. benchmark.json retains the Hamiltonian, configuration, selected parameters, complete optimization and confirmation records, per-basis counts and seeds, exact checks and aggregate statistics. Both language reports are generated from that same benchmark result.
{
"artifacts": {
"data": "artifacts/finite-shot-vqe/benchmark.json",
"report_en": "artifacts/finite-shot-vqe/vqe-en.html",
"report_zh": "artifacts/finite-shot-vqe/vqe-zh.html"
},
"ground_energy": -0.706225774829855,
"objective_evaluation_ceiling": 12,
"paired_repeats": 3,
"sdk_version": "1.0.8a",
"strategies": [
{
"estimator_error_rmse": null,
"mean_exact_selected_energy": -0.12351605330345887,
"mean_paired_exact_reference_gap": 0.0,
"strategy": "exact"
},
{
"estimator_error_rmse": 0.01341777011131602,
"mean_exact_selected_energy": -0.1163455933580984,
"mean_paired_exact_reference_gap": 0.007170459945360484,
"strategy": "sampled_single"
},
{
"estimator_error_rmse": 0.03697706494314943,
"mean_exact_selected_energy": -0.11853708251671276,
"mean_paired_exact_reference_gap": 0.004978970786746112,
"strategy": "sampled_fixed"
},
{
"estimator_error_rmse": 0.04777476599368023,
"mean_exact_selected_energy": -0.11556891288183073,
"mean_paired_exact_reference_gap": 0.00794714042162815,
"strategy": "sampled_adaptive"
}
],
"total_backend_executions": 229,
"total_shots": 26624
}
In the default development run, all twelve strategy runs together used 229 Backend executions and 26,624 shots. Those include all four cost categories. The mean exact energy at selected parameters was about −0.12352 for the exact strategy and −0.11635 for the single-measurement strategy, both well above the ground energy. This small configuration demonstrates the experiment procedure; it has not converged to the optimum.
Read three quantities separately:
selected_estimator_energyis the independent confirmation estimate for a sampled strategy. It is not its lowest historical sample.estimator_errorsubtracts the exact energy at those same selected parameters. It measures estimation error at that point.paired_exact_reference_gapsubtracts the selected energy of the exact optimization run in the same repetition. It compares parameter quality with that particular optimization run, not with the true ground state.
A negative paired gap means the sampled strategy happened to find better parameters in that pair; it does not prove it is generally better. All nine sampled selections in the default run are inconclusive: confirmation still returns a selected candidate, but does not resolve a confident separation from its competitor.
Across the three repetitions, the benchmark reports sample variance, standard error and a two-sided Student-t interval for the mean. With only three observations, the 95% critical value is about 4.303, so intervals can be wide. These intervals describe variation across repeated runs. Within-candidate measurement errors answer a different question. Neither kind of interval proves a global optimum, and the most favorable estimate among several strategies is itself subject to selection effects.
"""Compare exact and sampled VQE objectives under a paired evaluation budget."""
from __future__ import annotations
import argparse
import json
from math import sqrt
from pathlib import Path
from typing import Any, Literal
import cascaqit
from cascaqit import (
VQE,
HamiltonianTerm,
OptimizerConfig,
PauliHamiltonian,
PauliMeasurementConfig,
PauliX,
PauliZ,
SampledSelectionConfig,
SPSAConfig,
VQESamplingBenchmarkConfig,
)
def experiment(
output_dir: Path, *, repeats: int = 3, shots: int = 128, budget: int = 12
) -> dict[str, Any]:
"""Reuse executed benchmark data for both language reports and the JSON file."""
hamiltonian = PauliHamiltonian(
hamiltonian_id="project.finite-shot.hamiltonian",
logical_order=("q0",),
constant=0.1,
terms=(
HamiltonianTerm("x", 0.7, PauliX("q0")),
HamiltonianTerm("z", -0.4, PauliZ("q0")),
),
)
vqe = VQE(hamiltonian, algorithm_id="project.finite-shot.vqe")
config = VQESamplingBenchmarkConfig(
repeats=repeats,
objective_evaluation_budget=budget,
optimizer=OptimizerConfig(
method="SPSA",
max_iterations=1,
spsa=SPSAConfig(learning_rate=0.2, perturbation=0.15),
),
measurement=PauliMeasurementConfig(
shots_per_group=shots, allocation="coefficient_l1"
),
sampled_selection=SampledSelectionConfig(
candidate_count=2, repeats_per_candidate=2
),
fixed_objective_repeats=2,
adaptive_objective_repeats=2,
adaptive_max_objective_repeats=3,
adaptive_standard_error_target=0.08,
final_shots=shots,
confidence_level=0.95,
root_seed=29,
)
result = vqe.benchmark_sampling(config)
output_dir.mkdir(parents=True, exist_ok=True)
artifacts = {"data": str(output_dir / "benchmark.json")}
languages: tuple[Literal["en", "zh"], ...] = ("en", "zh")
for language in languages:
path = output_dir / f"vqe-{language}.html"
result.report(
path,
language=language,
title="Finite-shot VQE" if language == "en" else "有限采样 VQE",
)
artifacts[f"report_{language}"] = str(path)
rows = []
for run in result.runs:
selection = run.result.sampled_selection
rows.append(
{
"strategy": run.strategy,
"repeat_index": run.repeat_index,
"paired_seed": run.paired_seed,
"selected_parameters": dict(
run.result.selected_evaluation.parameter_bind.values
),
"selected_estimator_energy": run.selected_estimator_energy,
"exact_selected_energy": run.exact_selected_energy,
"estimator_error": run.estimator_error,
"paired_exact_reference_gap": run.paired_exact_reference_gap,
"objective_evaluations": run.objective_evaluation_count,
"objective_backend_executions": run.objective_backend_execution_count,
"objective_shots": run.objective_shots,
"confirmation_backend_executions": (
run.confirmation_backend_execution_count
),
"confirmation_shots": run.confirmation_shots,
"final_sampling_backend_executions": (
run.final_sampling_backend_execution_count
),
"final_sampling_shots": run.final_sampling_shots,
"diagnostic_exact_backend_executions": (
run.diagnostic_exact_backend_execution_count
),
"selection_status": None if selection is None else selection.status,
}
)
payload = {
"sdk_version": cascaqit.__version__,
"hamiltonian": hamiltonian.to_dict(),
"ground_energy": 0.1 - sqrt(0.65),
"config": config.to_dict(),
"rows": rows,
"benchmark": result.to_dict(),
"artifacts": artifacts,
}
Path(artifacts["data"]).write_text(
json.dumps(payload, indent=2) + "\n", encoding="utf-8"
)
return {
"sdk_version": cascaqit.__version__,
"paired_repeats": repeats,
"objective_evaluation_ceiling": budget,
"ground_energy": payload["ground_energy"],
"strategies": [
{
"strategy": item.strategy,
"mean_exact_selected_energy": item.mean_exact_selected_energy,
"mean_paired_exact_reference_gap": item.mean_paired_exact_reference_gap,
"estimator_error_rmse": item.estimator_error_rmse,
}
for item in result.statistics
],
"total_backend_executions": sum(
row[key]
for row in rows
for key in (
"objective_backend_executions",
"confirmation_backend_executions",
"final_sampling_backend_executions",
"diagnostic_exact_backend_executions",
)
),
"total_shots": sum(
row[key]
for row in rows
for key in ("objective_shots", "confirmation_shots", "final_sampling_shots")
),
"artifacts": artifacts,
}
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--output-dir", type=Path, default=Path("artifacts/finite-shot-vqe")
)
parser.add_argument("--repeats", type=int, default=3)
parser.add_argument("--shots", type=int, default=128)
parser.add_argument("--budget", type=int, default=12)
args = parser.parse_args()
print(
json.dumps(
experiment(
args.output_dir,
repeats=args.repeats,
shots=args.shots,
budget=args.budget,
),
sort_keys=True,
)
)
if __name__ == "__main__":
main()
Submit the comparison with its costs¶
Save the script revision, dependency versions, JSON and report. Include all repetitions in your comparison, even if a strategy performs poorly.
- Use the analytic energy formula to check each selected parameter set. Compare both its paired-reference gap and its gap to
E0. - Reconstruct one sampled energy directly from its X/Z counts. Explain why the final Z counts alone are insufficient.
- Run
--shots 512 --output-dir artifacts/vqe-more-shots. Compare within-evaluation standard errors and actual Backend/shot costs. Does every selected energy improve? - Run
--budget 16 --output-dir artifacts/vqe-more-budget. The iteration cap remains one. Explain why raising only the ceiling may leave actual usage unchanged. - In a copy, set
max_iterations=3and choose a budget that covers the repeated evaluations. Then run--repeats 10. Compare paired differences and their intervals before recommending a strategy.
Check your reasoning
The exact energy must lie above E0, while a sampled estimate can fall below it through statistical error. Reconstruct X and Z means as (n0−n1)/N in their respective bases, then apply the coefficients and constant. Increasing shots usually reduces measurement uncertainty at a fixed state, but adaptive stopping and different parameter updates can change the complete experiment. It does not guarantee an improved selected energy in every run. An unused ceiling provides no extra iterations. With three SPSA updates, initial/plus/minus/final estimates require up to eight logical points: fixed repetition can require 16 complete measurements, while adaptive repetition can require 24. Choose the budget before inspecting outcomes.
For a research extension, compare strategies at a fixed total shot budget that includes confirmation, specify stopping rules in advance, and use enough repetitions to characterize the observed variation. Keep the exact diagnostic results separate in the cost discussion. This one-qubit test provides no evidence of quantum advantage or hardware performance.