LATENT REFERENCES / TAG2
SARSA
Original title: SARSA
This reference note belongs to Tag2 in Latent References, an archive curated by Keigo Yoshida. Its archive region is Physics. The note preserves its source text and links so that readers can trace the material behind the 3D map.
- Collection
- Tag2
- Archive region
- Physics
Archived reference note
English translation of the archived note. JP shows the original text. Source links and literal code are retained; the translation does not update or independently verify the source claims.
Sarsa. Sarsa is a learning method consisting of 5 elements: “S (current state),” “A (agent's action),” “R (reward),” “S' (state after the action),” and “A' (the agent's next action determined from the post-action state).” When an agent takes an action from its current state, it receives a reward for that action.
https://scrapbox.io/files/64c656f603f8d3001c4209c1.png
Sarsa. Sarsaとは、「S(現在の状態)」「A(エージェントの行動)」「R(報酬)」「S'(行動後の状態)」「A'(行動後の状態から判断した、エージェントの次の行動)」の5つの要素から構成される学習方法です。 現在の状態からエージェントがある行動を取ったとき、エージェントには行動に対する報酬が与えられます。
https://scrapbox.io/files/64c656f603f8d3001c4209c1.png
Source updated 2023-07-30 · Snapshot 2026-10-08
Source links and calculated neighbors
Cosine values measure shared lexical features, not truth, agreement or identical meaning. Original reference links are labeled separately.
- Swarm IntelligenceComputed lexical cosine similarity 0.115 · shared title, text, tags and references
- GridworldComputed lexical cosine similarity 0.094 · shared title, text, tags and references
- MCP Model Context ProtocolComputed lexical cosine similarity 0.086 · shared title, text, tags and references