基于两阶段分层DQN的电动汽车实时充电调度策略
DOI:
作者:
作者单位:

华东交通大学

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学(72161011)


Real-Time Charging Scheduling Strategy for Electric Vehicles Based on a Two-Stage Hierarchical DQN
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对电动汽车充电路径规划中现有方法忽略排队随机性且难以兼顾实时性与全局最优的问题,提出一种基于广义总成本(TGC)与分层深度Q网络(DQN)的路径优化方法。首先,引入M/M/c排队模型量化充电等待时间的随机性,通过TGC体系协助深度Q网络将非线性排队、行驶时间与分时电价统一为单一标量奖励。其次,结合Top-M预筛选机制与动作有效性掩码来解决大规模动作空间导致的收敛难题,并利用经验回放打破数据相关性。实验在构建的高保真仿真环境中进行性能评估,在Top-M参数设为10时,该策略的全局最优解命中率达89.5%,平均相对后悔值收敛至0.0122,单次推理耗时缩短至0.35ms。

    Abstract:

    To address the issues in existing electric vehicle charging path planning methods, namely the neglect of queuing randomness and the difficulty in balancing real-time performance with global optimality, this paper proposes a path optimization method based on Generalized Total Cost (TGC) and a hierarchical Deep Q-Network (DQN). First, an M/M/c queuing model is introduced to quantify the randomness of charging waiting time. The TGC framework assists the Deep Q-Network in unifying nonlinear queuing time, travel time, and time-of-use electricity price into a single scalar reward. Second, a Top-M pre-screening mechanism combined with an action validity mask is employed to resolve convergence difficulties caused by large action spaces, and experience replay is utilized to break data correlations. The method""s performance is evaluated in a constructed high-fidelity simulation environment. When the Top-M parameter is set to 10, the proposed strategy achieves a global optimal solution hit rate of 89.5%, an average relative regret value converging to 0.0122, and a single inference time reduced to 0.35 ms.

    参考文献
    相似文献
    引证文献
引用本文
分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-05-11
  • 最后修改日期:2026-06-13
  • 录用日期:2026-06-22
  • 在线发布日期: 2026-07-23
  • 出版日期:
关闭