Open Access
ARTICLE
A Bilevel Deep Learning Optimization Framework for Joint Energy Harvesting Prediction and Energy-Aware Scheduling in IoT-Based Wireless Sensor Networks
1 Department of Renewable Energy, Jadara University, Irbid, Jordan
2 Department of Cybersecurity, Irbid National University, Irbid, Jordan
3 Department of Computer Science, College of Information Technology, Amman Arab University, Amman, Jordan
4 Faculty of Information Technology, Al-Ahliyya Amman University, Amman, Jordan
5 Department of Information Technology, Faculty of Prince Al-Hussien bin Abdullah, The Hashemite University, Zarqa, Jordan
6 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
7 Department of Computer Science and Artificial Intelligence, Umm Al-Qura University, Makkah, Saudi Arabia
* Corresponding Author: Bashar S. Khassawneh. Email:
Computers, Materials & Continua 2026, 88(3), 38 https://doi.org/10.32604/cmc.2026.079984
Received 01 February 2026; Accepted 22 April 2026; Issue published 23 July 2026
Abstract
Energy sustainability and secure operation are persistent challenges in Internet-of-Things (IoT) wireless sensor networks (WSNs), where limited battery capacity, heterogeneous traffic, and security procedures jointly drive premature node depletion and service degradation. This paper proposes an uncertainty-aware bilevel co-optimization framework that unifies residual-energy prediction with robust, energy-aware scheduling for clustered IoT-WSNs. At the lower level, a lightweight temporal predictor (TCN + LSTM with stochastic sampling) learns short-horizon residual-energy evolution from multivariate, dataset-aligned windows capturing sensing/communication activity, proximity-to-cluster-head effects, and security overhead (authentication latency, key exchange, and rekeying), and produces both point forecasts and uncertainty estimates to enable risk-sensitive control. At the upper level, a constrained, horizon-based scheduler selects per-node actions (duty cycle, sensing rate, transmission power) to extend network lifetime and balance residual energy while enforcing safety thresholds and operational bounds; bilevel coupling is realized via differentiable hypergradient updates, complemented by trust-region action smoothing and adaptive primal–dual constraint handling to suppress energy-critical states under uncertainty. On a real-world WSN energy–security dataset, the proposed model attains the best lower-level learning performance withKeywords
The pervasive integration of Internet of Things (IoT) sensing infrastructure into cyber-physical environments has elevated wireless sensor networks (WSNs) from simple data collectors to critical components of intelligent monitoring and control in smart cities, industrial automation, environmental surveillance, and resilient infrastructure [1]. In such deployments, energy sustainability is a first-order operational constraint: once nodes deplete their batteries, the network experiences coverage gaps, degraded connectivity, increased packet loss, and ultimately service disruption. In many modern WSN paradigms, particularly energy-harvesting-enabled systems, node operation is influenced not only by energy consumption but also by intermittent and environment-dependent energy replenishment processes (solar or ambient harvesting). In practical deployments, the dominant and directly observable quantity remains the residual battery energy, which reflects the net effect of both energy consumption and any implicit or unobserved energy inflows. Unlike idealized models, real WSN energy consumption is shaped by multiple coupled processes, including sensing and local processing, wireless transmission and reception, cluster coordination, and security procedures such as authentication and cryptographic key management. These factors interact nonlinearly and vary over time due to heterogeneous node roles, clustering topology, proximity-dependent radio costs, and dynamic security overheads, producing non-stationary energy trajectories that are difficult to manage using static scheduling rules [2]. Accordingly, modeling residual energy evolution provides a control-relevant abstraction of system dynamics, even in scenarios where explicit energy harvesting processes are not directly measured.
Recent progress in data-driven modeling and deep learning has enabled more expressive representations of WSN dynamics, particularly for temporal prediction tasks. Many existing methods treat residual energy prediction and scheduling as separable stages: a predictor is trained offline to minimize point-wise error, while a scheduler is designed independently, often using deterministic assumptions that treat predictions as [3]. In the context of energy-harvesting WSNs, several works attempt to incorporate harvesting-aware models; however, these often rely on simplified or assumed energy inflow distributions, limiting their applicability under real-world uncertainty and heterogeneous operating conditions. This decoupling is problematic in practice because the cost of error is asymmetric and constraint-sensitive. Overestimating residual energy can lead to aggressive duty cycles, sensing rates, or transmit power settings that accelerate depletion and trigger early node failure, whereas underestimation primarily reduces throughput or sensing fidelity. Moreover, network operation is inherently multi-objective and constrained, requiring simultaneous consideration of lifetime maximization, energy balance, communication performance, and security compliance. Consequently, predictive accuracy alone is insufficient; what is needed is control-relevant prediction that is uncertainty-aware and explicitly aligned with long-term operational objectives [4]. This is particularly important when residual energy is used as a proxy for system sustainability in the absence of explicit harvesting measurements.
To address these challenges, this paper proposes a unified deep learning-based bilevel co-optimization framework that tightly couples residual energy prediction (lower level) with energy-aware scheduling (upper level) in a closed-loop learning–control architecture. Rather than explicitly modeling stochastic energy harvesting processes, the proposed framework focuses on learning the temporal evolution of residual energy as a measurable and control-relevant system state, capturing the net effect of workload, communication, and security-induced energy consumption. At the lower level, a lightweight temporal model combining temporal convolution and gated recurrent processing learns node-level energy evolution from multivariate observations that capture energy states, packet statistics, clustering context, and security-induced overhead. In addition to point forecasts, uncertainty is estimated via stochastic inference to quantify confidence under non-stationary consumption regimes. At the upper level, scheduling is formulated as a constrained robust optimization problem that adapts duty cycle, sensing rate, and transmission power to reduce energy depletion and imbalance while enforcing explicit safety thresholds and feasible action bounds. The coupling between the two levels is realized through hypergradient-based updates, allowing the long-term scheduling objective to shape the learning process and thereby producing a predictor optimized for downstream sustainability rather than short-horizon regression error. This design ensures that the framework remains applicable to both conventional battery-powered WSNs and future extensions involving explicit energy harvesting models.
The proposed framework offers many contributions as follows:
• Uncertainty-aware energy prediction: A lightweight TCN + LSTM model predicts node residual energy and estimates uncertainty using energy, traffic, clustering, and security-overhead features.
• Bilevel robust scheduling: A constrained upper-level optimizer adapts duty cycle, sensing rate, and transmit power to improve lifetime and energy balance under safety and security costs.
• Differentiable closed-loop co-optimization: Hypergradient coupling with trust-region and primal–dual updates enables online learning–control adaptation via replay-buffer refinement.
Research on energy-aware wireless sensor networks (WSNs) has evolved along several largely separate directions, including protocol-driven scheduling and routing, data-driven prediction of energy or network states, and learning-based control or policy optimization. Although each line of work contributes useful insights, the literature remains fragmented in terms of how prediction and decision-making are integrated. In particular, most existing studies either (i) optimize scheduling using hand-crafted or heuristic rules without a learned predictive model, (ii) perform prediction as an isolated task without explicitly linking it to downstream control objectives, or (iii) learn control policies directly but rely on fixed state abstractions, simplified dynamics, or weak safety treatment. This fragmentation creates a methodological gap in the design of control-aware predictive systems for WSNs, especially under uncertainty, security overhead, and heterogeneous node behavior. The present work addresses this gap by formulating residual-energy prediction and scheduling as a hierarchically coupled bilevel optimization problem in which the predictive model is explicitly shaped by downstream control requirements.
Sah et al. [5]. This paper introduces ERAS, an efficient routing awareness-based scheduling scheme for solar energy-harvesting WSNs that integrates hierarchical clustering with TDMA synchronization to curb redundant transmissions and extend lifetime. The study targets the coupled problems of cluster-head (CH) election and route discovery under variable harvested energy, defining the problem as minimizing CH count subject to coverage and energy-balance constraints while modeling radio (free-space/multipath) and solar EH dynamics. Methodologically, ERAS selects CHs using a composite score of residual and harvested energy, schedules CM activity via TDMA sleep/awake control, aggregates data at CHs, and uses single-/multi-hop inter-cluster routing; complexity and message-exchange bounds are derived, and MATLAB simulations compare ERAS with ERA, SEED, TM2RP, and CREST on two network sizes (100–200 nodes). ERAS achieves the best lifetime and stability: lifetime share
Oudah and Shaker [6]. This survey reviews Wireless Sensor Networks (WSNs) for renewable energy monitoring, covering architectures, deployments (PV, wind, hydro, microgrids), and energy-aware routing (LEACH/PEGASIS/TEEN/ZRP and recent variants). The objective is to map protocol capabilities to domain constraints and identify performance bottlenecks in latency, lifetime, and reliability. Methodologically, it synthesizes protocol taxonomies, application case studies, and emerging enablers (edge/AI, blockchain, federated learning, 6G). Key results include protocol–use-case matching (OLSR for frequent transmissions; AODV for event-driven monitoring) and quantified gains from optimized clustering; EE-OLEACH reports
Alrayes et al. [7]. The paper proposes a privacy-preserving statistical learning pipeline for IoT intrusion detection (PPSLOA-HDBDE) that targets high-dimensional big-data settings by combining linear scaling normalization (LSN), Sand Cat Swarm Optimization (SCSO) for feature selection, an ensemble of TCN/MAE/XGBoost classifiers, and hyperparameter tuning via an Improved Marine Predator Algorithm (IMPA). The stated objective is to preserve data confidentiality while maintaining analytical efficacy; the problem is framed as an accurate multi-class IDS under dimensionality, privacy, and computational constraints. The system normalizes inputs, selects features with SCSO (fitness balances accuracy and subset size), trains the TCN/MAE/XGBoost ensemble, and optimizes with IMPA (Brownian/Levy-phase search plus opposition-based learning). On BoT-IoT, the approach reports best performance at 1500 epochs: Accuracy
Fereidunian et al. [8]. This paper tackles online power control in large energy-harvesting (EH) wireless networks by learning policies with a deep neural network (DNN), addressing the intractability of stochastic control/MDP formulations under high-dimensional state spaces. The objective is to approximate the optimal online transmit-power mapping that maximizes time-averaged throughput without noncausal information. The authors (i) solve many offline convex instances (with noncausal EH/channel knowledge) to generate optimal state
Table 1 clarifies the central gap in the existing literature. Prior studies generally focus on one of three isolated perspectives: protocol-centric scheduling, prediction-only modeling, or learning-based control. While each perspective addresses an important part of the broader problem, none simultaneously provides: (i) a temporally grounded and uncertainty-aware predictor of node energy evolution, (ii) a constrained scheduler that explicitly accounts for safety, efficiency, and security overheads, and (iii) a bilevel coupling mechanism through which the downstream control objective can directly shape the learning process. This missing integration is particularly significant in practical IoT-WSNs, where energy dynamics are non-stationary, operational constraints are strict, and purely decoupled pipelines can lead to suboptimal or unsafe control behavior. The proposed framework is designed precisely to close this gap by moving beyond sequential prediction-then-control and toward a unified control-aware learning architecture.

The rapid proliferation of Internet of Things (IoT)-based wireless sensor networks (WSNs) has intensified the challenge of maintaining long-term, secure, and energy-efficient operation under highly constrained node resources and dynamic environmental conditions. In practical deployments, node energy depletion is influenced not only by sensing and communication workloads but also by security mechanisms such as authentication and key management, which introduce non-negligible and time-varying overheads. Traditional energy-aware scheduling approaches typically rely on static heuristics or offline-trained predictors, limiting their ability to adapt to non-stationary energy dynamics and uncertainty. To address these limitations, this paper proposes a unified deep learning–driven bilevel optimization framework that jointly learns residual energy evolution and optimizes scheduling decisions in a closed-loop manner. The overall methodology is illustrated in Fig. 1, which depicts the complete workflow from dataset preparation and preprocessing to uncertainty-aware residual energy prediction, and finally to bilevel energy-aware scheduling and optimization. As shown in the figure, encoded and normalized WSN data are used to train a lower-level deep predictor that estimates future residual energy along with uncertainty. At the same time, an upper-level controller leverages these predictions to perform robust, constraint-aware scheduling using safety filters and multi-objective criteria. Hypergradient-based coupling enables feedback from the scheduling objective to refine the predictor, and an online replay buffer supports continual adaptation under evolving network conditions. This tightly integrated learning–control loop forms the foundation of the proposed methodology, enabling sustainable, secure, and adaptive energy management for IoT-based WSNs. Throughout this section, we denote by

Figure 1: Methodology framework.
The experiments in this study utilize a real-world Wireless Sensor Network (WSN) dataset [9], which captures detailed node-level information related to energy dynamics, communication behavior, and security operations in clustered IoT environments. The dataset comprises approximately 50,000 records collected from

3.2 System Model and Energy Dynamics
We consider an IoT-based wireless sensor network consisting of a set of sensor nodes
Each node starts with an initial energy budget
The residual energy state of each node is recorded in residual_energy and evolves dynamically in response a both energy consumption and potential energy replenishment processes. In practical WSN deployments, nodes may harvest energy from environmental sources such as solar radiation or ambient energy. Although the dataset does not explicitly include a dedicated harvesting variable, the observed residual energy implicitly reflects the net effect of energy consumption and energy replenishment over time.
Nodes with excessive depletion may become non-functional and/or compromised, as indicated by compromised_status. Network sustainability is evaluated through the dataset-provided network_lifetime indicator and the global notion of time-to-first-node-death, both of which quantify whether secure and energy-efficient operation can be maintained over prolonged periods [12]. This system model provides the physical foundation for the proposed bilevel framework, in which the lower level learns energy-evolution patterns and the upper level optimizes scheduling and role-management decisions to extend network lifetime while preserving security.
Let
where
Communication energy is expressed as a function of proximity and packet activity [14]:
where
Security overhead is modeled as:
where
A node is operational if
Table 3 summarizes the principal state variables and dataset-aligned parameters used to model node energy evolution and network behavior, with explicit clarification of their computational roles to ensure transparency and prevent methodological ambiguity. In particular, the table distinguishes between input features, target variables, derived quantities, and evaluation-only metrics, thereby addressing potential concerns regarding data leakage and reproducibility. Residual energy

Algorithm 1 formalizes the dataset-aligned energy evolution process as a causal, time-consistent update mechanism that integrates communication, sensing, security, and latent environmental effects within a unified energy balance framework. At each discrete time step

3.3 Modeling Scope: Residual Energy vs. Energy Harvesting Processes
Energy dynamics in wireless sensor networks (WSNs) can be formally described as the interaction between energy consumption processes-driven by sensing, communication, and security operations-and energy replenishment processes arising from environmental harvesting sources such as solar or ambient energy. In principle, accurate modeling of energy-harvesting WSNs requires explicit characterization of both components; however, harvesting processes are inherently stochastic, environment-dependent, and often only partially observable in real-world deployments. Many practical datasets, including the one used in this study, do not provide direct measurements of harvested energy, but instead report the resulting residual battery state. From a system-theoretic and control-oriented perspective, residual energy can therefore be interpreted as a sufficient and observable state variable that implicitly captures the net effect of both consumption and any latent energy inflow processes. To ensure conceptual consistency among the problem formulation, dataset characteristics, and evaluation protocol, this work adopts a residual-energy-centric modeling strategy, in which learning and control are performed directly on the measurable energy state rather than on an explicitly decomposed harvesting model. This design choice resolves the apparent mismatch between energy-harvesting terminology and residual-energy-based implementation, while preserving applicability to realistic WSN deployments where harvesting signals are unavailable or unreliable.
The energy evolution of node
Table 4 provides a structured, multidimensional comparison between explicit energy-harvesting formulations and the proposed residual-energy-centric modeling paradigm, highlighting key differences in system representation, observability, modeling assumptions, and control integration. As shown in the table, conventional approaches rely on a dual-state formulation

To further clarify the methodological novelty, the proposed framework can be formally distinguished from three major classes of existing approaches. First, in classical model predictive control (MPC), system dynamics are assumed to be known or separately estimated, and optimization is performed using fixed predictive models. In contrast, the proposed framework jointly optimizes the predictive model and control policy via bilevel coupling, thereby enabling the control objectives to directly influence the learned energy dynamics. Second, reinforcement learning (RL)-based scheduling methods typically learn policies through trial-and-error interaction with the environment, without explicit modeling of system dynamics or safety guarantees. The proposed approach instead maintains an explicit predictive model with uncertainty quantification and enforces constraint-aware optimization, resulting in more stable and interpretable decision-making. Third, meta-learning and adaptive control methods focus on rapid parameter adaptation but do not explicitly incorporate hierarchical optimization between prediction and control. The proposed bilevel formulation explicitly captures this hierarchical dependency through hypergradient-based coupling, ensuring that the predictor is optimized for downstream control performance rather than standalone prediction accuracy.
Fig. 2 provides a comprehensive conceptual and architectural clarification of the proposed residual- energy-centric modeling paradigm, explicitly distinguishing it from traditional energy-harvesting-based formulations and from the approach adopted in this work. As illustrated in panel (A), conventional models rely on a decomposed representation of energy dynamics, where the system state is defined by both battery energy

Figure 2: Residual-energy-centric modeling as a sufficient observable state.
3.4 Deep Learning Model for Energy Prediction
The lower-level component of the proposed bilevel framework is a deep learning model that learns the temporal evolution of node energy states from historical observations. The model takes as input a multivariate time-series window constructed from dataset-aligned features, including residual energy, cumulative energy usage, packet transmission and reception counts, proximity to the cluster head, and a security-related overhead indicator. Given a fixed input window of length
The architecture follows a lightweight temporal modeling design suitable for IoT-scale data, combining a temporal convolutional layer with gated recurrent processing. Temporal Convolutional Network (TCN) layers extract local and long-range temporal patterns via dilated causal convolutions, while recurrent layers capture sequential dependencies across sensing rounds. To improve robustness to non-stationary energy consumption, normalization and dropout layers are incorporated to ensure stable training across heterogeneous node behavior. The final fully connected layer maps the learned temporal representation to a scalar prediction of future residual energy for each node.
To train the model, a regression-based objective is used to minimize the discrepancy between predicted and observed residual energy values. In addition to the primary prediction loss, regularization terms are introduced to penalize unstable temporal fluctuation and reduce overfitting. The trained model outputs both a point prediction and an uncertainty-aware confidence measure, which are later used by the upper-level optimization to avoid aggressive scheduling decisions when prediction uncertainty is high. These deep learning components serve as the adaptive predictive engine that continuously informs energy-aware control policy in the bilevel framework.
Mathematical Formulation
Let
This formulation allows the model to predict future energy states based on recent system behavior, enabling proactive scheduling decisions at the upper level.
The training objective minimizes the mean squared prediction error with regularization:
where
Table 5 summarizes the configuration of the deep learning model used for energy prediction. The combination of temporal convolutional and recurrent layers enables the model to efficiently capture both short-term fluctuations and long-term trends in node energy evolution. lightweight design choices ensure compatibility with a resource-constrained IoT environment, while dropout layers support uncertainty-aware inference. These architectures provide accurate and stable energy predictions, which directly influence the effectiveness of the upper-level scheduling optimization in Algorithm 2.


In simple terms, the model learns how each node’s energy level evolves by analyzing recent historical observations. The temporal convolutional layers capture short-term patterns, while the recurrent layer models longer-term dependencies across sensing rounds.
3.5 Operational Definitions, Robust MPC Objective, and Feasibility Guarantees
The upper-level scheduling formulation is grounded in an explicit mapping between control actions and measurable physical quantities, ensuring that each optimization term is operationally defined, auditable, and directly computable from the available data. For each node

Uncertainty in future energy states is modeled through the predictive distribution provided by the lower-level model, where residual energy follows
3.6 Energy-Aware Scheduling and Bilevel Optimization Formulation
The upper level of the proposed framework focuses on energy-aware scheduling decisions that regulate node activity to maximize long-term network sustainability under dynamic and uncertain operating conditions. The scheduling variables include the duty cycle, sensing rate, and transmission power of each node, which jointly determine the operational workload and energy expenditure. Specifically, the duty cycle
The scheduling objective is defined at the network level and aims to extend operational lifetime while ensuring secure and efficient communication [17]. Network sustainability is quantified through both expected energy consumption and the variance of residual energy across nodes, reflecting the need to avoid energy imbalance and early node failure. At each scheduling round, the controller selects actions that minimize long-term energy usage while promoting fairness in energy distribution across heterogeneous node roles. The formulation explicitly accounts for security-induced energy overhead and node heterogeneity, resulting in a multi-objective, constraint-aware optimization problem. At the lower level, a deep learning model parameterized by
subject to operational constraints:
The lower-level learning problem optimizes the deep model parameters
The bilevel optimization problem is defined as:
constraints in (13).
Unlike conventional predict-then-control pipelines, where the predictive model is trained independently and treated as fixed during control optimization, the proposed bilevel formulation introduces an explicit dependency between control variables and the optimal predictor parameters through the mapping
At optimality, the lower-level parameters satisfy the first-order stationarity condition:
Differentiating this condition with respect to the upper-level variables
Substituting into the total derivative of the upper-level objective gives:
where the vector
with
Fig. 3 presents the proposed Energy-Aware Scheduling and Bilevel Co-Optimization Framework, which couples an upper-level scheduling controller with a lower-level residual-energy predictor in a closed-loop learning-control cycle. At the lower level, node histories are organized into an input window and processed by the predictor

Figure 3: Energy-aware scheduling and bilevel co-optimization framework.
Table 7 summarizes the principal decision variables, algorithmic components, and computational mechanisms involved in the bilevel optimization process. In particular, it explicitly characterizes how hypergradients are computed through implicit differentiation, how second-order information is incorporated via Hessian–vector products, and how numerical stability is ensured through damping and trust-region constraints. To improve practical relevance, the table also clarifies execution locality (sensor node, cluster head, or edge controller), indicative runtime burden, memory implications, and scalability trends with respect to the number of nodes and temporal horizon. These additions show that the extra cost introduced by bilevel coupling remains concentrated in edge-side computation and scales in a controlled manner relative to standard learning-based scheduling pipelines.

The computational complexity of the hypergradient update is determined by solving a linear system involving the Hessian of the lower-level loss. Instead of explicitly forming the Hessian matrix, which would incur
Algorithm 3 describes the proposed bilevel co-optimization procedure as a closed-loop learning–control process that iteratively couples uncertainty-aware prediction with robust scheduling. At each time step, historical node observations are organized into temporal windows and processed by the predictor

This section presents the experimental evaluation of the proposed uncertainty-aware bilevel optimization framework for energy-efficient scheduling in IoT-based wireless sensor networks. The results are organized to assess both the predictive capability of the lower-level deep learning model and the system-level impact of the upper-level scheduling policy. First, the accuracy and robustness of residual energy prediction are examined using standard regression and classification metrics, highlighting the effect of temporal modeling and uncertainty estimation. Next, the performance of the bilevel scheduling mechanism is analyzed with respect to network lifetime, energy balance, satisfaction of safety constraints, and operational overhead, and is compared with representative baseline strategies. Ablation analyses are used to isolate the contribution of individual components-such as uncertainty-aware prediction, trust-region control, and hypergradient coupling-to overall system performance. Together, these results demonstrate that improved prediction quality translates into safer, more sustainable scheduling decisions under realistic, security-aware WSN operating conditions.
4.1 Experimental Setup and Reproducibility Protocol
To ensure rigorous and reproducible evaluation, the proposed framework is assessed through a controlled simulation environment based on the real-world WSN dataset described in Section 3.1. The experimental setup explicitly defines network configuration, temporal dynamics, and model training conditions. The simulated wireless sensor network consists of
For the upper-level scheduling, control variables include duty cycle (
Table 8 presents the quantitative evaluation of the proposed and baseline models using averaged results over multiple independent runs. Conventional regression models capture a substantial portion of the variance in residual energy, with linear and ridge baselines achieving

Tree-based ensemble methods improve performance by modeling nonlinear interactions, reducing MAE to
The proposed TCN + LSTM model demonstrates a clear and consistent advantage, achieving MAE of
For the classification task, the proposed model achieves an F1-score of
Table 9 presents the quantitative scheduling performance of different control strategies, averaged over multiple simulation runs. The results reveal a clear hierarchy, where progressively tighter integration between prediction and control yields significant improvements in network sustainability. Static scheduling serves as a baseline but suffers from inefficient energy utilization because it cannot adapt to dynamic workloads and security overhead. Greedy energy-aware policies provide marginal improvements; however, their lack of uncertainty modeling leads to unstable behavior and inconsistent reductions in violations under stochastic conditions.

Incorporating prediction through a DL-only strategy improves both network lifetime and energy balance, as the controller becomes proactive rather than reactive. Nevertheless, without bilevel coupling, the predictor remains optimized for point-wise accuracy rather than control feasibility, limiting overall system performance. The ablation analysis highlights the contribution of each component. Removing the uncertainty-aware risk envelope significantly increases the number of safety violations (
In Fig. 4A,B (transient energy responses), the fixed policy exhibits the least stable behavior: in Fig. 4A it undergoes a deep undershoot to approximately

Figure 4: Discussion of transient energy response: (A) transient response around the normalized target level near 1.0 under fixed policy, DL-only, and the full proposed method; (B) transient response around the lower normalized operating level near 0.70 under the same control strategies.
In Fig. 5A of the utility/throughput tracking results (oracle vs. fixed policy), the oracle starts near

Figure 5: Discussion of utility/throughput tracking performance. (A) Tracking comparison between the oracle reference and the fixed policy. (B) Tracking comparison between the oracle reference and the DL-only strategy without bilevel coupling. (C) Tracking comparison between the oracle reference and the full proposed algorithm.
In Fig. 5B,C (DL-only vs. full proposed tracking), the DL-only strategy substantially improves fidelity but still shows noticeable error during abrupt transitions: following the first drop at
This paper presented a unified deep learning–driven bilevel optimization framework for energy-aware scheduling in IoT-based wireless sensor networks. Unlike conventional approaches that decouple energy prediction from control decisions, the proposed methodology jointly optimizes a residual energy predictor and a constrained scheduling policy within a closed-loop learning–control architecture. By explicitly modeling temporal energy dynamics, security-induced overheads, and prediction uncertainty, the framework enables adaptive and risk-aware scheduling that is robust to non-stationary network conditions. At the lower level, a lightweight TCN–LSTM predictor learns node-level residual energy evolution from real WSN data while providing uncertainty estimates that quantify forecast reliability. These predictions are not used in isolation; instead, they directly inform the upper-level controller, which optimizes duty cycle, sensing rate, and transmission power under explicit safety constraints, trust-region bounds, and multi-objective sustainability criteria. Hypergradient-based bilevel coupling ensures that the learning process is shaped by long-term network objectives rather than short-term prediction accuracy alone, transforming the predictor into a control-aware component of the system. Analytical evaluation and expected-performance analysis indicate that the proposed framework can substantially reduce energy prediction error while improving scheduling safety, energy balance, and network lifetime compared to static or single-level baselines. Importantly, the framework’s modular design and reliance on lightweight models make it suitable for deployment in resource-constrained IoT environments. Future work will extend this framework to energy-harvesting WSNs, distributed and federated learning settings, and real-time hardware deployments. In addition, incorporating adaptive security policies and multi-hop routing decisions represents a promising direction for further enhancing the sustainable and secure operation of IoT networks.
Acknowledgement: The authors extend their appreciation to Umm Al-Qura University, Saudi Arabia for funding this research work through grant number: 26UQU4270203GSSR01.
Funding Statement: This research work was funded by Umm Al-Qura University, Saudi Arabia, under Grant Number: 26UQU4270203GSSR01.
Author Contributions: Mohammad Q. Al-Jamal: conceptualization; data curation; formal analysis; investigation; methodology; software; validation; visualization; writing—original draft. Mahmoud Al Jamal: data curation; formal analysis; investigation; methodology; software; validation; visualization; writing—original draft. Bashar S. Khassawneh: conceptualization; methodology; supervision; writing—review & editing. Ayoub Alsarhan: formal analysis; investigation; methodology; validation; writing—review & editing. Amina Salhi: validation; resources; supervision; writing—review & editing. Tahani Alsubait: validation; resources; supervision; writing—review & editing. All authors reviewed and approved the final version of the manuscript.
Availability of Data and Materials: The data supporting the findings of this study are publicly available on Kaggle under the title “WSN Dataset for Secure and Energy Efficient” at: https://www.kaggle.com/datasets/ziya07/wsn-dataset-for-secure-and-energy-efficient (accessed on 10 April 2026).
Ethics Approval: Not applicable.
Conflicts of Interest: The authors declare no conflicts of interest.
References
1. Sharma MK, Zappone A, Debbah M, Assaad M. Deep learning based online power control for large energy harvesting networks. In: Proceedings of the ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2019 May 12–17; Brighton, UK. p. 8429–33. [Google Scholar]
2. Iyobhebhe M, Tekanyi AMS, Abubilal KA, Usman AD, Isiaku Y, Agbon EE, et al. A review on energy consumption model on hierarchical clustering techniques for IoT-based multilevel heterogeneous WSNs using energy aware node selection. Vokasi Unesa Bull Eng Technol Appl Sci. 2025;2(2):100–11. doi:10.26740/vubeta.v2i2.34882. [Google Scholar] [CrossRef]
3. Goswami P, Khan T, Pathak V, Alabdultif A. Machine learning based dynamic trust estimation framework for securing wireless sensor networks. Sci Rep. 2025;15(1):35821. doi:10.1038/s41598-025-19768-z. [Google Scholar] [PubMed] [CrossRef]
4. Tawfeek MA, Alrashdi I, Alruwaili M, Talaat FM. A fuzzy multi-objective framework for energy optimization and reliable routing in wireless sensor networks via particle swarm optimization. Comput Mater Contin. 2025;83(2):2773–92. doi:10.32604/cmc.2025.061773. [Google Scholar] [CrossRef]
5. Sah DK, Hazra A, Mazumdar N, Amgoth T. An efficient routing awareness based scheduling approach in energy harvesting wireless sensor networks. IEEE Sens J. 2023;23(15):17638–47. doi:10.1109/JSEN.2023.3279249. [Google Scholar] [CrossRef]
6. Oudah AY, Shaker LM. Wireless sensor networks in renewable energy monitoring: a comprehensive survey of protocols and applications. AUIQ Tech Eng Sci. 2025;2(2):8. doi:10.70645/3078-3437.1037. [Google Scholar] [CrossRef]
7. Alrayes FS, Maray M, Alshuhail A, Almustafa KM, Darem AA, Al-Sharafi AM, et al. Privacy-preserving approach for IoT networks using statistical learning with optimization algorithm on high-dimensional big data environment. Sci Rep. 2025;15(1):3338. doi:10.1038/s41598-025-87454-1. [Google Scholar] [PubMed] [CrossRef]
8. Fereidunian A, Alimoradi Z, Sebt MA. Energy smart environments: toward a systems perspective framework for the energy dimension of smart systems of systems. In: Proceedings of the 2025 6th International Conference on Optimizing Electrical Energy Consumption (OEEC); 2025 Feb 25–26; Najafabad, Iran. p. 1–10. doi:10.1109/OEEC66525.2025.11100143. [Google Scholar] [CrossRef]
9. Ziya07. WSN dataset for secure and energy efficient wireless sensor networks. [cited 2026 Jan 29]. Available from: https://www.kaggle.com/datasets/ziya07/wsn-dataset-for-secure-and-energy-efficient. [Google Scholar]
10. Zeng W, Wang Z, Yang S, He D, Chan S. ICMH-CHR: an intra-cluster multi-hop based cluster head rotation protocol for wireless sensor networks. Ad Hoc Netw. 2025;173(4):103829. doi:10.1016/j.adhoc.2025.103829. [Google Scholar] [CrossRef]
11. Zhang Y, Yang L, Tan Y. Energy-efficient adaptive routing in heterogeneous wireless sensor networks via hybrid PSO and dynamic clustering. J Cloud Comput. 2025;14(1):46. doi:10.1186/s13677-025-00768-3. [Google Scholar] [CrossRef]
12. Akram M, Ullah Bazai S, Imran Ghafoor M, Akram S, Ilyas QM, Mehmood A, et al. EEMLCR: energy-efficient machine learning-based clustering and routing for wireless sensor networks. IEEE Access. 2025;13(17):70849–71. doi:10.1109/ACCESS.2025.3562368. [Google Scholar] [CrossRef]
13. Naderi MY, Basagni S, Chowdhury KR. Modeling the residual energy and lifetime of energy harvesting sensor nodes. In: Proceedings of the 2012 IEEE Global Communications Conference (GLOBECOM); 2012 Dec 3–7; Anaheim, CA, USA. p. 3394–400. doi:10.1109/GLOCOM.2012.6503639. [Google Scholar] [CrossRef]
14. Freschi V, Lattanzi E. A study on the impact of packet length on communication in low power wireless sensor networks under interference. IEEE Internet Things J. 2019;6(2):3820–30. doi:10.1109/JIOT.2019.2891841. [Google Scholar] [CrossRef]
15. Tonin A, Nasser Ali Algabri M, Ali MN, Ghurab M, Ahmed Ghaleb Al-Khulaidi A, Ahmed Al-Baltah I. A novel energy model for optimizing energy efficiency and prolonging network lifetime in MANET. IEEE Access. 2025;13(3):117771–87. doi:10.1109/ACCESS.2025.3586795. [Google Scholar] [CrossRef]
16. Alanazi R, Obayya M, Alghamdi AM, Nemri N, Alshahrani S, Alduaiji N, et al. Machine learning-driven routing optimization for energy-efficient 6G-enabled wireless sensor networks. Alex Eng J. 2025;129(1):877–88. doi:10.1016/j.aej.2025.07.032. [Google Scholar] [CrossRef]
17. Singh A, Raj A, Rani P, Khatibi A, Aldeeb H, Shukla PK, et al. Resilient wireless sensor networks in industrial contexts via energy-efficient optimization and trust-based secure routing. Peer Peer Netw Appl. 2025;18(3):132. doi:10.1007/s12083-025-01946-5. [Google Scholar] [CrossRef]
18. Kaddi M, Omari M, Salameh K, Alnoman A. Energy-efficient clustering in wireless sensor networks using grey wolf optimization and enhanced CSMA/CA. Sensors. 2024;24(16):5234. doi:10.3390/s24165234. [Google Scholar] [PubMed] [CrossRef]
19. Jiang L, Xiao Q, Chen T. Improved analysis of penalty-based methods for bilevel optimization with coupled constraints. In: Proceedings of the 2025 33rd European Signal Processing Conference (EUSIPCO); 2025 Sep 8–12; Palermo, Italy. p. 1183–7. doi:10.23919/EUSIPCO63237.2025.11226224. [Google Scholar] [CrossRef]
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools