DQL still learning at evaluation time

Issue forked from #87 by @kvas7andy

[learner.epsilon_greedy_search(...)](https://github.com/microsoft/CyberBattleSim/blob/4fd228bccfc2b088d911e27072a923251203cac8/cyberbattle/agents/baseline/learner.py#L126) is often used for training agents with different algorithms, including DQL in the `dql_run`. However `dql_exploit_run` with input network `dql_run` as policy-agent and `eval_episode_count` parameter for the number of episodes, gives an impression that runs are used for evaluation of the trained DQN. The only distinguishable difference between 2 runs is epsilon queal to 0, which leads to exploitation mode of training, but does not exclude training, because during run with [ learner.epsilon_greedy_search](https://github.com/microsoft/CyberBattleSim/blob/4fd228bccfc2b088d911e27072a923251203cac8/cyberbattle/agents/baseline/learner.py#L126) the `optimizer.step()` is executed on each step of training in the file `agent_dql.py`, function call [learner.on_step(...)](https://github.com/microsoft/CyberBattleSim/blob/4fd228bccfc2b088d911e27072a923251203cac8/cyberbattle/agents/baseline/agent_dql.py#L348).
- **Solution**: I will include in Pull request the code I used for better evaluation (based on  [ learner.epsilon_greedy_search(...)](https://github.com/microsoft/CyberBattleSim/blob/4fd228bccfc2b088d911e27072a923251203cac8/cyberbattle/agents/baseline/learner.py#L126) and generate pictures below. 
- **Screenshots**: Figure 1 & 2 and figure 3 & 4 , shows result of chain network evaluation using corresponding new cell in [notebook_benchmark-chain.ipynb](https://github.com/microsoft/CyberBattleSim/blob/4fd228bccfc2b088d911e27072a923251203cac8/notebooks/notebook_benchmark-chain.ipynb). As you can see on [figure 1](https://user-images.githubusercontent.com/8929593/194348810-c5731ab6-80bd-4fd3-af8f-e2070e0aa943.png)  training on the initial 50 episodes is not enough for owning 100% of the network (AttackerGoal), whereas original run `dql_exploit_run` internally using `learner.on_step(...)` [figure 2](https://user-images.githubusercontent.com/8929593/194348854-9569a9cc-f553-48ec-b352-4ffc0890bf40.png) leads to much better results, due to optimization process, which still process ongoing experience of agent. We can overcome this inaccurate evaluation and still reach the goal in 100% of times [figure 3](https://user-images.githubusercontent.com/8929593/194349289-349268d9-3e2c-47d3-a6b8-6e0c602bfba0.png), while training on 200 episodes with commented `learner.on_step()`. It fixes trained network and stops optimizing during evaluation, but leads to the ownership of all the network with larger amount of learning episodes. This means with 200 episodes it is feasible to learn optimal path of agent attacks inside chain network configuration. 
Lastly, [figure 4](https://user-images.githubusercontent.com/8929593/194349308-ad6e719c-a5e8-4174-8b7d-e3ca0b71b358.png) we can compare those runs with correct evaluation runs on 20 episodes reach 6000+ and 120+  cumulative reward for for 200 and 50 training episodes correspondently.
[Figure 1: (after PR) no optimizer during evaluation, 20 trained episodes, 20 evaluation episodes](https://user-images.githubusercontent.com/8929593/194348810-c5731ab6-80bd-4fd3-af8f-e2070e0aa943.png)
[Figure 2: (before  & after PR) dql_exploit_run with optimizer during evaluation, 20 trained episodes, 5 evaluation episodes](https://user-images.githubusercontent.com/8929593/194348854-9569a9cc-f553-48ec-b352-4ffc0890bf40.png)
[Figure3: (after PR) no optimizer during evaluation, **200** trained episodes, 20 evaluation episodes](https://user-images.githubusercontent.com/8929593/194349289-349268d9-3e2c-47d3-a6b8-6e0c602bfba0.png)
[Figure 4: (after PR) comparison of evaluation for network trained on 200 and 20 episodes, chain network configuration](https://user-images.githubusercontent.com/8929593/194349308-ad6e719c-a5e8-4174-8b7d-e3ca0b71b358.png)

[ToyCTF](https://github.com/microsoft/CyberBattleSim/blob/4fd228bccfc2b088d911e27072a923251203cac8/notebooks/notebook_benchmark-toyctf.ipynb) benchmark is inaccurate, because with correct evaluation procedure, like with chain network configuration, agent does not reqch goal of 6 owned nodes after 200 training episodes.


Provide feedback

Saved searches

Use saved searches to filter your results more quickly

DQL still learning at evaluation time #115

Metadata

Assignees

Labels

Type

Fields

Projects

Milestone

Relationships

Development

DQL still learning at evaluation time #115

Description

Metadata

Metadata

Assignees

Labels

Type

Fields

Projects

Milestone

Relationships

Development

Issue actions