Skip to content

About

Physics-accurate free-space quantum network simulator for LEO satellite entanglement distribution — models photon loss, atmospheric attenuation, fidelity degradation, iterative purification, and quantum memory decoherence with real orbital mechanics

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PPO-Based Entanglement Scheduler for LEO Satellite Quantum Networks

Python PyTorch Stable-Baselines3 License: CC0--1.0 Citation


Overview

This repository presents a physics-grounded reinforcement-learning framework for scheduling entanglement requests in Low Earth Orbit (LEO) satellite quantum networks. It formulates the scheduling problem as a constrained Markov Decision Process (MDP) and solves it with Proximal Policy Optimization (PPO) and dynamic action masking.

The simulator combines satellite-visibility information with a physical-layer model of free-space diffraction, atmospheric attenuation, link-fidelity degradation, and iterative entanglement purification. At every decision step, the policy selects from the five highest-fidelity candidate satellites or defers a request when the available links cannot satisfy the operational constraints.

This public implementation demonstrates the environment design, RL formulation, and benchmarking workflow developed for the project. An extended research manuscript is in preparation.


Key Contributions

  • Physics-grounded network simulation for free-space satellite quantum links, including optical loss, atmospheric attenuation, fidelity estimation, and purification.
  • Constrained PPO scheduler that jointly considers request service, link fidelity, queue expiration, visibility, and exclusive satellite allocation.
  • Dynamic action masking that prevents invalid satellite selections caused by missing simultaneous line-of-sight or an allocation already made during the current macro-step.
  • Micro-step scheduling that evaluates active requests sequentially inside a 22 ms physical macro-step, reducing the decision space while preserving time consistency.
  • Benchmarking against a fidelity-greedy scheduler under identical visibility and queue conditions.

Mathematical Formulation

Key Variables

Symbol Description
$N_{\mathrm{served},t}$ Requests completed at time step $t$
$N_{\mathrm{active},t}$ Active requests evaluated during the macro-step
$F_{\mathrm{avg},t}$ Average post-purification fidelity of served links at time step $t$
$N_{\mathrm{expired},t}$ Requests dropped because their waiting time exceeded the TTL
$\alpha, \beta, \delta$ Weights for service, fidelity, and expiry penalty
$\pi$ Satellite-selection policy
TTL Request time-to-live / coherence-window setting

Objective

The scheduler maximizes request service and link quality while penalizing expired requests:

$\max_{\pi} \sum_{t=0}^{T} \left(\alpha \frac{N_{\mathrm{served},t}(\pi)}{N_{\mathrm{active},t}} + \beta F_{\mathrm{avg},t} - \delta N_{\mathrm{expired},t}(\pi) \right).$

The expression inside the summation is used as the PPO reward. The default implementation uses $\alpha=0.45$, $\beta=0.30$, and $\delta=0.25$.

MDP Definition

State $S_t \in \mathbb{R}^{59}$ combines the active request and the five highest-fidelity candidate satellites:

  • Request features (14 dimensions): source ground station, destination ground station, normalized wait time, and generated-pair count.
  • Satellite features (45 dimensions): elevation, slant range, expected fidelity, and a six-bit line-of-sight coverage mask for each candidate satellite.

Action $A_t \in {0,\ldots,5}$ selects one of the five candidate satellites or a defer action. The action mask removes satellites that do not simultaneously cover the selected source-destination pair or have already been allocated in the current macro-step.


Results

PPO versus Fidelity-Greedy Scheduling

Across TTL settings from 1 to 5 seconds, PPO maintains approximately 110--120 requests/s, while the greedy scheduler declines from approximately 65 requests/s at TTL = 1 s to approximately 49 requests/s at TTL = 5 s.

PPO also maintains a substantially lower number of expired requests throughout the TTL sweep.

Throughput versus TTL

Throughput comparison across TTL = 1--5 seconds.

Expired requests versus TTL

Expired-request comparison across TTL = 1--5 seconds.

Representative Detailed Evaluation

For the representative configuration below, the best PPO reward weighting ($\alpha=0.30$, $\beta=0.45$, $\delta=0.25$) achieved 120 requests/s, compared with 65 requests/s for the fidelity-greedy baseline: an 84.6% throughput improvement. It served 86,705 of 86,888 generated requests (a 99.79% completion rate) and recorded 183 expired requests, compared with 1,163 for the baseline: 6.35$\times$ fewer expirations.

Agent $\alpha$ $\beta$ $\delta$ Throughput (requests/s) Requests Generated Requests Served Completion Rate Expired
PPO 0.45 0.30 0.25 93 67,875 67,491 99.43% 384
PPO 0.30 0.45 0.25 120 86,888 86,705 99.79% 183
PPO 0.45 0.25 0.30 111 80,186 79,909 99.66% 277
PPO 0.25 0.30 0.45 114 82,645 82,420 99.73% 225
Greedy --- --- --- 65 46,920 45,757 97.52% 1,163

Repository Structure

PPO3V3/                         Dynamic-queue PPO implementation
├── quantum_env.py               Custom Gymnasium environment and action-mask logic
├── physics_engine.py            Optical-link and fidelity model
├── train_ppo.py                 PPO training workflow
├── test_ppo.py                  Policy evaluation workflow
├── SimulatorGreedy_alignedToPPO3.py
│                                 Fidelity-greedy baseline
└── plot_results.py              PPO-versus-greedy visualisation

PPO3/                            Earlier implementation
images/                          Benchmark figures
CITATION.Cff                     Citation metadata
LICENSE                          CC0 1.0 Universal

Citation

If this work is useful in your research, please cite the repository using CITATION.Cff.

License

This project is released under the CC0 1.0 Universal public-domain dedication.

Contact

Muhammad Tauseef Mushtaq PhD, Department of Electrical and Information Engineering Politecnico di Bari, Italy 📧 m.mushtaq@phd.poliba.it 🔗 https://www.linkedin.com/in/tauseef-mushtaq/

About

Physics-accurate free-space quantum network simulator for LEO satellite entanglement distribution — models photon loss, atmospheric attenuation, fidelity degradation, iterative purification, and quantum memory decoherence with real orbital mechanics

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages