This repository presents a physics-grounded reinforcement-learning framework for scheduling entanglement requests in Low Earth Orbit (LEO) satellite quantum networks. It formulates the scheduling problem as a constrained Markov Decision Process (MDP) and solves it with Proximal Policy Optimization (PPO) and dynamic action masking.
The simulator combines satellite-visibility information with a physical-layer model of free-space diffraction, atmospheric attenuation, link-fidelity degradation, and iterative entanglement purification. At every decision step, the policy selects from the five highest-fidelity candidate satellites or defers a request when the available links cannot satisfy the operational constraints.
This public implementation demonstrates the environment design, RL formulation, and benchmarking workflow developed for the project. An extended research manuscript is in preparation.
- Physics-grounded network simulation for free-space satellite quantum links, including optical loss, atmospheric attenuation, fidelity estimation, and purification.
- Constrained PPO scheduler that jointly considers request service, link fidelity, queue expiration, visibility, and exclusive satellite allocation.
- Dynamic action masking that prevents invalid satellite selections caused by missing simultaneous line-of-sight or an allocation already made during the current macro-step.
- Micro-step scheduling that evaluates active requests sequentially inside a 22 ms physical macro-step, reducing the decision space while preserving time consistency.
- Benchmarking against a fidelity-greedy scheduler under identical visibility and queue conditions.
| Symbol | Description |
|---|---|
| Requests completed at time step |
|
| Active requests evaluated during the macro-step | |
| Average post-purification fidelity of served links at time step |
|
| Requests dropped because their waiting time exceeded the TTL | |
| Weights for service, fidelity, and expiry penalty | |
| Satellite-selection policy | |
| TTL | Request time-to-live / coherence-window setting |
The scheduler maximizes request service and link quality while penalizing expired requests:
The expression inside the summation is used as the PPO reward. The default implementation uses
State
- Request features (14 dimensions): source ground station, destination ground station, normalized wait time, and generated-pair count.
- Satellite features (45 dimensions): elevation, slant range, expected fidelity, and a six-bit line-of-sight coverage mask for each candidate satellite.
Action
Across TTL settings from 1 to 5 seconds, PPO maintains approximately 110--120 requests/s, while the greedy scheduler declines from approximately 65 requests/s at TTL = 1 s to approximately 49 requests/s at TTL = 5 s.
PPO also maintains a substantially lower number of expired requests throughout the TTL sweep.
Throughput comparison across TTL = 1--5 seconds.
Expired-request comparison across TTL = 1--5 seconds.
For the representative configuration below, the best PPO reward weighting (
| Agent | Throughput (requests/s) | Requests Generated | Requests Served | Completion Rate | Expired | |||
|---|---|---|---|---|---|---|---|---|
| PPO | 0.45 | 0.30 | 0.25 | 93 | 67,875 | 67,491 | 99.43% | 384 |
| PPO | 0.30 | 0.45 | 0.25 | 120 | 86,888 | 86,705 | 99.79% | 183 |
| PPO | 0.45 | 0.25 | 0.30 | 111 | 80,186 | 79,909 | 99.66% | 277 |
| PPO | 0.25 | 0.30 | 0.45 | 114 | 82,645 | 82,420 | 99.73% | 225 |
| Greedy | --- | --- | --- | 65 | 46,920 | 45,757 | 97.52% | 1,163 |
PPO3V3/ Dynamic-queue PPO implementation
├── quantum_env.py Custom Gymnasium environment and action-mask logic
├── physics_engine.py Optical-link and fidelity model
├── train_ppo.py PPO training workflow
├── test_ppo.py Policy evaluation workflow
├── SimulatorGreedy_alignedToPPO3.py
│ Fidelity-greedy baseline
└── plot_results.py PPO-versus-greedy visualisation
PPO3/ Earlier implementation
images/ Benchmark figures
CITATION.Cff Citation metadata
LICENSE CC0 1.0 Universal
If this work is useful in your research, please cite the repository using CITATION.Cff.
This project is released under the CC0 1.0 Universal public-domain dedication.
Muhammad Tauseef Mushtaq PhD, Department of Electrical and Information Engineering Politecnico di Bari, Italy 📧 m.mushtaq@phd.poliba.it 🔗 https://www.linkedin.com/in/tauseef-mushtaq/

