Use Microsoft 365 Copilot as an LLM backend for OpenAI-compatible coding agents like pi and OpenClaw. Wraps M365 Copilot's WebSocket/SignalR API in an OpenAI-compatible interface with tool calling support.
Want the gory protocol details? See docs/m365-copilot-api.md — a full write-up of M365 Copilot's undocumented WebSocket API: auth, SignalR frames, tones/models, throttling, the "Disengaged" filter, and the Copilot Studio agent trick that makes tool calling work.
M365 Copilot uses a SignalR WebSocket protocol, not the OpenAI API. This project translates between the two:
- Standalone proxy — HTTP server with
/v1/chat/completionsand/v1/modelsendpoints. Works with any OpenAI-compatible client (pi, OpenClaw, etc.). - OpenClaw plugin — Config generator + setup CLI for OpenClaw's provider system.
M365 Copilot doesn't support OpenAI-style tool_calls natively. Instead, tools are
emulated via a Markdown-fence format (the JSON {"tool":...} format was removed —
it scored 0/5 on real agentic tasks; see hypotheses §9):
- Tool definitions are injected into the prompt as fenced templates inside a
<tools>block - The model emits a fenced tool call — a code block whose info-string is the tool name
(scalar args as
key: valueheader lines, one free-form body arg as the fence body,old/newedits as aider-styleSEARCH/REPLACEdiffs) - The proxy/handler parses that and converts it to OpenAI
tool_callsformat - Shell-routing (the key lever): M365's chat-tuned model won't "act as an agent" on
demand but will reflexively write a
```bashblock. When the toolset includes a shell tool (bash/shell/run/run_command/… — any name), the proxy injects "do the whole step by writing one```bashblock" framing and routes that block to the shell tool. This exploits the one agentic behavior Microsoft's system prompt permits, and is what turns 0/5 into real multi-turn loops (verified 9-tool-call bug fix). - Reliability comes from the Copilot Studio agent (below) + the fenced/shell framing — without the agent, M365 ignores tool instructions and answers in prose
On first use, the system creates a Copilot Studio agent with tool-calling instructions baked into its server-side system prompt. This is done via the PowerPlatform API:
- Discovers the environment URL via the BAP API (
api.bap.microsoft.com) - Creates a bot with instructions in the Copilot Studio
minimalBotsAPI - Publishes the bot to get a
TitleId - Uses the agent ID (
T_{titleId}.{botId}.gpt.default) in WebSocket chat requests - Caches the agent ID in
~/.config/opencode-m365/agent-id.json
Each agent session reuses the same M365 conversation (same sessionId + conversationId). The WebSocket reconnects per turn but M365 maintains server-side context. This saves quota — the 600 message limit applies per-conversation.
@m365-copilot/core — Shared: auth, WebSocket client, tool formatting, proxy server, agent management, session
├── @m365-copilot/proxy — Standalone HTTP proxy binary
└── @m365-copilot/openclaw-plugin — OpenClaw config generator + setup CLI + skill
- Node.js 24+
- pnpm 10+
- An M365 account with Copilot access
- A way to sign in, either:
- TOTP-based MFA with the base32 secret in hand — the automated login types the 6-digit code itself, so it needs the seed, not an app on your phone. See Getting the TOTP secret. This is the headless-friendly option; prefer it if your tenant allows it.
- or a display and 30 seconds — set
M365_ENABLE_INTERACTIVE_APPROVAL=1and sign in by hand once. Required if your tenant has no TOTP option (push-only MFA, FIDO2, Okta/Ping/Duo).
git clone https://github.com/cramt/m365-copilot-proxy
cd m365-copilot-proxy
pnpm install
pnpm buildCreate ~/.config/opencode-m365/secrets.json:
{
"email": "you@company.com",
"password": "your-password",
"mfaSecret": "YOUR_TOTP_BASE32_SECRET"
}mfaSecret is the base32 seed your authenticator app derives its 6-digit codes
from — not a code itself. It looks like JBSWY3DPEHPK3PXP: 16–32 characters, A–Z
and 2–7 only, no spaces.
Already using a password manager for TOTP? Just read it back out. 1Password,
Bitwarden, KeePassXC, Aegis, Ente Auth and friends all keep the seed and will show it
on demand — open the item's one-time-password field and reveal it. You'll get either a
bare base32 string or an otpauth://totp/...?secret=JBSWY3DPEHPK3PXP&... URI; in the
second case the secret= parameter is the bit you want. No re-enrollment needed.
If the seed is trapped in Microsoft Authenticator, enroll a second method. That app deliberately never exposes it, and Microsoft's security-info page won't re-display the key after enrollment either — between them that's the only case where the seed is genuinely unrecoverable. Enroll a fresh entry and copy it on the way past:
- Go to https://aka.ms/mfasetup (My Account → Security info).
- Add sign-in method → Authenticator app.
- Click "I want to use a different authenticator app". This step is the one that matters — the default path assumes Microsoft Authenticator and registers a push-only method with no seed you can extract.
- At the QR-code screen, click "Can't scan image?". It reveals a Secret key — that's your base32 string.
- Generate a code from it to finish enrollment —
oathtool --totp -b <secret>, or paste the seed into your password manager — and enter that code.
Store it somewhere that will give it back (see above) so this is a one-time chore. The new entry sits alongside your existing sign-in methods; you don't have to remove Microsoft Authenticator.
Plenty of tenants can't do the above, and then there is no seed to extract:
- the authenticator-app / software-OATH method is disabled by tenant policy;
- MFA is push / number-matching only, FIDO2, Windows Hello, or certificate-based;
- sign-in is federated to a third-party IdP (Okta, Ping, Duo) that owns the MFA step.
The stored-credentials flow above cannot work in those cases — the automated login has no code to type, and no amount of configuration fixes that.
Use interactive approval instead. Skip secrets.json entirely and set:
export M365_ENABLE_INTERACTIVE_APPROVAL=1
m365-proxy 4141A real browser window opens; you complete SSO/MFA by hand exactly as you would in Outlook — push, FIDO2, Okta, whatever your tenant enforces. The proxy captures the resulting OAuth code and from then on refreshes tokens silently from the MSAL cache, so the window is a one-time cost, not a per-run prompt. It also kicks in when stored credentials exist but stop working (policy change, MFA method swap).
Contributed by @EatonWu. Two honest caveats:
- It is opt-in on purpose. Without the flag a headless host (systemd, CI, a second
PC) fails loudly instead of hanging on a window nobody can see. Set
M365_NO_INTERACTIVE=1to veto it outright. - The redirect capture and token exchange are the same ones the automated path uses in production, but nobody has yet run this end to end against a federated Okta/Ping/Duo tenant. If that's you, please report in #4 — working or not.
Not device code. The obvious headless alternative —
login.microsoft.com/devicewith a short code — is permanently dead here, not merely unimplemented. Initiation returns a valid code, but redeeming it fails withAADSTS7000218: Entra treats Microsoft's own Copilot client as confidential for that grant and demands aclient_secretonly Microsoft holds. No tenant admin can grant it. Measured, with the sign-in actually completed, in §13 H13.2.
On first run, the system does an automated browser login (via Playwright/Chromium) to get OAuth tokens. After that, tokens refresh silently from the MSAL cache.
Start the proxy:
m365-proxy 4143 # or: pnpm run proxy 4143Point pi at it via ~/.pi/agent/models.json:
{
"providers": {
"m365": {
"baseUrl": "http://localhost:4143/v1",
"api": "openai-completions",
"apiKey": "not-needed",
"compat": {
"supportsDeveloperRole": false,
"supportsReasoningEffort": false,
"supportsUsageInStreaming": false
},
"models": [
{ "id": "gpt-5.5-think-deeper", "name": "M365 Copilot (GPT-5.5, recommended)" },
{ "id": "m365-copilot", "name": "M365 Copilot (Auto)" }
]
}
}
}
apiKeyis a placeholder. The proxy binds to localhost and validates no credential — the field exists only because the OpenAI client schema requires it.
Then run pi (use gpt-5.5-think-deeper — the reliable tool-calling model — and keep the
toolset lean; M365 "disengages" on very large tool payloads, see
docs/m365-copilot-api.md):
pi --models "gpt-5.5-think-deeper" -p --tools read,list,edit,write "your task"This is verified working end-to-end, including multi-tool calls and real file edits.
# Configure and start in one command
m365-openclaw-setup --start
# Or configure only, then start separately
m365-openclaw-setup
m365-proxy 4141The proxy uses session reuse and delta messages — follow-up turns only send new messages, saving M365 quota. New conversations are detected automatically when the message array shrinks or the first user message changes.
npx m365-proxy 4141
# or
pnpm run devThen point any OpenAI-compatible client at http://localhost:4141/v1.
Proxy conversations use M365 temporary-chat mode by default
(disableMemory=1 on the authenticated Chathub URL), so they are not added to
the Copilot history sidebar while multi-turn context still works for the live
conversation. Set M365_SAVE_HISTORY=1 before starting the proxy only if you
deliberately want proxy chats saved in Copilot history.
The proxy is a Nitro service. The flake exposes a package (built from the workspace via pnpm2nix) and a NixOS module:
# flake.nix
{
inputs.m365.url = "github:cramt/m365-copilot-proxy";
outputs = { nixpkgs, m365, ... }: {
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
modules = [
m365.nixosModules.default
{
services.m365-copilot-proxy = {
enable = true;
# JSON with { email, password, mfaSecret } — kept out of the Nix store,
# delivered via systemd LoadCredential. Manage with sops-nix/agenix.
secretsFile = "/run/secrets/m365-copilot.json";
# port = 4141; # default
# host = "127.0.0.1"; # default — do not expose; unauthenticated, paid account
# openFirewall = false;
};
}
];
};
};
}The service runs as a hardened DynamicUser unit. Auth state (msal-cache.json,
agent-id.json) persists in /var/lib/m365-copilot-proxy; a fresh deploy self-bootstraps
via headless login using secretsFile + the bundled Chromium. To run the package directly
without NixOS: nix run github:cramt/m365-copilot-proxy -- 4141.
| Model ID | M365 Tone | Description |
|---|---|---|
gpt-6-think-deeper |
Gpt_6_Reasoning | GPT-6 reasoning. Needs a paid/premium Copilot seat (see below); 30/30 on the bench (agent-less, relay framing) |
gpt-5.6-think-deeper |
Gpt_5_6_Reasoning | GPT-5.6 reasoning — 27/30 on the bench, tied with gpt-5.5-think-deeper |
gpt-5.6 / gpt-5.6-quick |
Gpt_5_6_Chat | GPT-5.6 fast ("GPT 5.6 Quick response" in the web UI). Weak at tool calling — 7/30 on the bench (see below) |
gpt-5.5-think-deeper |
Gpt_5_5_Reasoning | Recommended default for agents/tool-calling — 26/30 on the bench |
gpt-5.5 / gpt-5.5-quick |
Gpt_5_5_Chat | GPT-5.5 fast — 16/30 on the bench, the strongest benchmarked chat tone |
m365-copilot / auto |
magic | Auto-routing — high-variance at tool-calling (confabulates; see below) |
quick |
Gpt_5_5_Chat | Alias of gpt-5.5 (its old Gpt_Quick tone was retired — see below) |
think-deeper |
Gpt_5_5_Reasoning | Alias of gpt-5.5-think-deeper (its old Gpt_Reasoning tone was retired) |
claude / claude-sonnet / claude-sonnet-4.6 |
Claude_Sonnet | Real Anthropic Claude Sonnet 4.6 (agent-less path) — 78/90 on the bench with the relay framing (47/76 with baseline). claude-sonnet-4.5 is kept as an alias |
claude-sonnet-5 |
Claude_Sonnet (paid scenario) | Claude Sonnet 5. Needs a paid/premium Copilot seat — same tone as above, the scenario picks the model. 27/30 on the bench with the relay framing (6/40 with baseline; see below) |
claude-sonnet-think-deeper |
Claude_Sonnet_Reasoning | Claude reasoning |
claude-opus / claude-opus-5 |
Claude_Opus | Real Claude Opus 5. Needs a paid/premium Copilot seat and has a small separate quota (see below) |
gpt-5.4 / gpt-5.4-quick |
Gpt_5_4_* | GPT-5.4 |
gpt-5.3 / gpt-5.3-think-deeper |
Gpt_5_3_* | GPT-5.3 |
gpt-5.2 / gpt-5.2-think-deeper |
Gpt_5_2_* | GPT-5.2 |
Bench scores are 10 tasks × 3 reps with the confab-retry off (M365_NO_CONFAB_RETRY=1), so
every first-try give-up counts, including ones the proxy's retry would normally recover
(hypotheses §19). GPT-6's 24/30 was measured with the retry on, so it
isn't directly comparable.
Microsoft retired the
*_Quicktones;*_Chattones replaced them. Everyquick/*-quickmodel ID still works and resolves to the replacement*_Chattone (e.g.gpt-5.3-quick→Gpt_5_3_Chat), so existing configs need no change.
✅ For tool calling, use
gpt-5.5-think-deeper(the default when no model is sent). The current agent + fenced/shell-routing path makes this reasoning tone robust: 26/30 on the bench, tied withgpt-5.6-think-deeperand well ahead of every benchmarked chat tone (docs/hypotheses.md §19), and 100% compliance and solve across prompt/toolset sizes in the tool-call harness (§12.10/§12.11). The defaultm365-copilot(magic) tone is not reliable for tools — it confabulates ("I no longer have access to the filesystem tools") and solves ~0% of real tasks (§12.11); a proxy request with nomodelfield already defaults togpt-5.5-think-deeperfor this reason.
⚠️ The older reasoning tones (gpt-5.2/gpt-5.3/gpt-5.4*-think-deeper, baregpt-5.4) route through M365'sDeepLeopipeline, which meta-analyzes the injected prompt and can disengage from tools. Prefergpt-5.5-think-deeper. See docs/m365-copilot-api.md §5/§10.
The web client calls this one "GPT 5.6 Quick response", which is the shape to expect:
it is the chat sibling of gpt-5.6-think-deeper, not a reasoning tone, and it self-identifies
as "the GPT-5 chat model" — exactly as Gpt_5_5_Chat does. Neither tone will tell you which
generation is actually answering, so don't pick between them on self-report.
Pick it for chat, not for tools. A higher version number is not an upgrade here: on the
bench it solves 7/30, against 16/30 for gpt-5.5. Its low message count (39 vs 69)
is not efficiency: 22 of its 23 failures gave up before making a single tool call. See
hypotheses §19.
No entitlement needed. It reaches a model on the default OfficeWebIncludedCopilot
scenario, so the proxy sends nothing special for it. This is newly true: until recently the
tone deflected with the canned BotConnection apology on the included scenario while serving
normally under OfficeWebPaidCopilot — it was entitlement-gated like Opus, which
hypotheses §12.15 had misfiled as a permanently dead route.
See §18 for the correction.
It needs the paid scenario, exactly like Opus. On the default
OfficeWebIncludedCopilot the Gpt_6_Reasoning tone is accepted and then deflects with the
canned apology; under OfficeWebPaidCopilot it serves normally. The proxy sends the paid
scenario automatically whenever the resolved tone is Gpt_6_Reasoning, so nothing has to be
configured — but this is an entitlement, not a bypass, and a seat without paid/premium
access will see the apology no matter which scenario is requested.
It is not metered separately, and that is the part worth stating. The paid scenario is an entitlement gate; it does not imply a budget. GPT-6 has no priority-access allowance and is throttled by the same per-conversation cap and thread-rate governor as every other model, so none of the Opus advice about rationing turns applies.
Tool calls go without the tool agent, and with the relay framing. With the Copilot Studio
tool agent attached, Gpt_6_Reasoning doesn't serve at all (a canned apology, on every account),
so the proxy sends GPT-6 tool requests agent-less, as it does Claude's (#41). Agent-less, under the
baseline framing GPT-6 worked in M365's own code interpreter instead of calling your tools:
0/30 on the bench. Under relay it scored 30/30 (confab-retry off, 2026-10-01), and 5/5 driving real pi. Both are the
defaults now; M365_FORCE_AGENT=1 and M365_FRAMING_VARIANT still override them.
gpt-5.5-think-deeper remains the recommended default and the no-model fallback, because it needs
no entitlement: making GPT-6 the default would hand most seats a model they can't reach.
Claude_Sonnet is Sonnet 4.6 on the included scenario and Sonnet 5 on the paid one, so
claude-sonnet-5 needs a paid/premium seat, like Opus. It is not separately metered.
Sonnet 5 arrives with its own tools (bash_tool, create_file, …) running in a remote
sandbox (/home/claude) that cannot see your files, and it reads the proxy's usual
<system>-tagged tool framing as a prompt injection — so with that framing it inspects its own
empty sandbox and reports that your files don't exist (6/40 on the bench). The proxy therefore
gives both Sonnet models a different default framing, relay: a plain note asking it to guide you through your
terminal one command at a time, which also tells it the sandbox is the wrong machine. That
scores 27/30 on the bench (from 5/30) and solved 5/5 fix-bug runs through real pi.
Sonnet 4.6 uses relay too (78/90 vs 47/76). Override with M365_FRAMING_VARIANT as usual.
Details: hypotheses §21.
The model behind this tone is Claude Opus 5. It performs very well here, with two things to know before you point an agent at it.
It needs the paid scenario. M365 gates the model list on the scenario sent with the
WebSocket connection. On the default OfficeWebIncludedCopilot the Opus tone is accepted but
never reaches a model — it returns a canned apology, which is why earlier notes in this repo
recorded Opus as a dead tone. The proxy now sends scenario=OfficeWebPaidCopilot (with the
licenseType=Premium that pairs with it) automatically whenever the resolved tone is
Claude_Opus; every model except gpt-6-think-deeper and claude-sonnet-5 keeps the included scenario. This is an
entitlement, not a bypass — your account has to actually hold paid/premium Copilot access, and licenseType
alone unlocks nothing. Override either with M365_SCENARIO / M365_LICENSE_TYPE.
It is metered separately, and the proxy spends it fast. Opus draws on a "priority access"
budget distinct from the ~600-message conversation cap. When it runs out, M365 replies with
text rather than an error — "You've used your available priority access to the Opus model for
today…" (or "…for the week"). Left alone that reads to an agent as the model's answer, so the
proxy detects both wordings and returns HTTP 429 (code: priority_access_exhausted) with a
Retry-After. Both budgets reset at midnight UTC; the weekly one on Monday.
Because agentic turns prepend a tool-framing block, driving Opus through a proxy burns that
budget faster than chatting with it by hand. Opus therefore defaults to the lean minimal
framing (~82% smaller than the default baseline, which exists to force M365's chat-tuned GPT
path to act and which Opus doesn't need). Whether the budget counts tokens or messages is
not verified — if it's per-message this saves latency rather than quota. Set
M365_FRAMING_VARIANT=baseline to opt out. See docs/hypotheses.md §15.
M365 Copilot generates images through a built-in server-side tool, and the core
package exposes it as one call. The picture comes back on a GraphicArt frame as
a URL (never as chat text), and the bytes sit behind a separate auth boundary —
generateImage() handles both, returning the image with bytes attached:
import { generateImage } from "@m365-copilot/core";
const [img] = await generateImage("A minimalist flat-design logo of a lighthouse, teal and white.");
// img.data -> Buffer (real PNG, verified end-to-end)
// img.base64 -> same bytes, ready for an OpenAI-style b64_json response
// img.contentType, img.size, img.orientationEverything the M365 web client can do is reachable — the proxy sends the same image optionsSets it does. Steer type and aspect with options (they nudge the prompt the way the GUI's meta-prompting does; the model still makes the final call):
await generateImage("a lighthouse on a cliff", { orientation: "portrait" }); // landscape | portrait | square
await generateImage("a lighthouse", { style: "icon" }); // natural | icon | story | designerYou don't have to call generateImage at all. Just like the web client, a plain
chat turn draws when asked — send "draw me an image of a green teapot" to
/v1/chat/completions (or ModelSession.run) with no tools and the image comes back
embedded in the reply as a markdown data-URI. (Image gen is enabled on the agent-less
path only, so it never competes with tool calling; set M365_NO_IMAGE_GEN=1 to force
pure text.)
Runs its own agent-less session. Uses the same login as chat; the artifact fetch uses
a designerappservice token acquired silently from the existing cache. Protocol
write-up: docs/hypotheses.md §14.
Separate, scarcer budget. Image generation draws on its own daily quota, distinct from the ~600-message conversation limit and not metered by the chat throttle — treat image calls as the expensive ones. When it's exhausted,
generateImage()throwsImageGenerationErrorwithreason: "quota_exceeded"(map it to HTTP 429); a plain chat turn instead returns M365's "can't generate any more images today" message as text.
An OpenAI-compatible POST /v1/images/generations endpoint on top of this is the
next step — the core API it needs is already in place.
The auth flow uses Azure MSAL with PKCE:
- Silent refresh — cached tokens from
~/.config/opencode-m365/msal-cache.json. The normal path; costs nothing and opens nothing. - Automated login — headless Playwright browser driving the AAD form with stored
credentials + a TOTP code generated from
mfaSecret. - Interactive approval — visible browser, human completes SSO/MFA (§13). Only when
M365_ENABLE_INTERACTIVE_APPROVAL=1, and only after step 2 is unavailable or has actually failed. This is the path for tenants where no TOTP seed exists.
All three redeem the code against https://login.microsoftonline.com/common/oauth2/nativeclient
with PKCE. That redirect isn't a stylistic choice: the client is Microsoft's own Copilot app
(the Sydney scopes are granted to no other), so nobody can register a loopback URI — a
generated http://localhost:<port> callback is rejected with AADSTS50011, and the device-code
grant demands a client_secret only Microsoft holds (AADSTS7000218). Both measured live in
§13; don't spend probes re-deriving them.
Three token scopes are acquired:
substrate.office.com/sydney/*— For M365 Copilot chatapi.powerplatform.com/.default— For Copilot Studio agent managementapi.bap.microsoft.com/.default— For environment discovery
| Variable | Description |
|---|---|
M365_DEBUG |
Set to 1 to enable debug logging to ~/.config/opencode-m365/debug.log (truncated payloads) |
M365_TRACE |
Set to 1 for full, untruncated debug logging (every WS frame/prompt/response) — implies M365_DEBUG. For reverse engineering. |
M365_LOG_STDOUT |
Set to 1 to mirror debug lines to the proxy's stdout as well as the log file, so you can watch a run without tailing it in a second terminal. Needs M365_DEBUG or M365_TRACE — on its own it logs nothing. |
M365_DUMP_FRAMES |
Set to 1 to write every WebSocket frame (both directions) to ~/.config/opencode-m365/frames/<requestId>.ndjson. For offline diffing of new M365 fields. |
M365_ALLOW_MULTI_TOOL |
Allow the model to emit multiple tool calls per turn (default: only the first is kept) |
M365_INJECT_REPLY_TOOL |
Set to 1 to inject a synthetic reply(text) tool. Forces every turn to be a tool call, including pure-prose answers. Cleaner contract for the model, +1 tool to the prompt (watch the Disengaged threshold). Confirmed 5/5 compliance on June 9 2026 (hypotheses §1.1). |
M365_NO_CONFAB_RETRY / M365_CONFAB_RETRIES |
M365's chat model sometimes produces prose instead of a tool call when it should act — either confabulating an inability ("I can't access the files, please paste them") or claiming a completion it never did ("I've replaced the README", with no tool call). By default the proxy detects both and re-prompts forcefully in the same conversation (M365_CONFAB_RETRIES, default 1) to force a real action. Set M365_NO_CONFAB_RETRY=1 to disable. |
M365_NO_BACKOFF (alias M365_NO_AUTO_REAUTH) |
Set to 1 to disable degradation backoff. By default, when empty/throttled responses span several distinct conversations in a short window (the thread-rate-throttle signature, F13), the proxy paces subsequent turns (a jittered delay before starting new backend conversations) to let the account self-heal. This replaced the old auto-reauth: a fresh login does not clear this throttle (it's oid-keyed — §11 H-R1) and raised our detection profile. A single long pi thread never trips the trigger. |
M365_BACKOFF_THRESHOLD / M365_BACKOFF_WINDOW_MS / M365_BACKOFF_BASE_MS / M365_BACKOFF_MAX_MS |
Tune backoff: distinct-conversation empties to trigger (default 3), the window they must fall in (default 120000), the initial pacing window (default 90000), and its escalation cap (default 600000). |
M365_BROWSER_PROFILE / M365_LOGIN_UA |
Override the persistent browser-profile dir and the login User-Agent used for the (rare) automated interactive login. The persistent profile keeps AAD SSO/device cookies so repeat logins are silent and look like a familiar device (§11 H-R3). |
M365_ENABLE_INTERACTIVE_APPROVAL |
Set to 1 to allow a visible browser window for sign-in when the automated login can't work or fails — the fallback for tenants with no TOTP option (push-only MFA, FIDO2, Okta/Ping/Duo). You complete SSO/MFA by hand once; tokens refresh silently afterwards. Off by default so headless hosts fail loudly rather than hang. See If your tenant has no TOTP option. |
M365_NO_INTERACTIVE |
Set to 1 to hard-disable any visible browser login, overriding the flag above. For systemd/CI hosts where a window must never open. |
M365_INTERACTIVE_TIMEOUT_MS |
How long to wait for you to finish the interactive sign-in (default 600000, i.e. 10 minutes). |
M365_LOGIN_LOCALE / M365_LOGIN_TIMEZONE |
Browser locale and timezone presented during login (defaults en-GB / Europe/Copenhagen). These are part of the anti-bot-scoring fingerprint (§11 F25) — set them to match your own machine if AAD starts treating your automated login as a bot. |
M365_SCENARIO / M365_LICENSE_TYPE |
Override the entitlement the WebSocket is opened under (defaults: OfficeWebIncludedCopilot / Starter, switching to OfficeWebPaidCopilot / Premium for the entitlement-gated models — claude-opus, gpt-6-think-deeper and claude-sonnet-5). scenario is what gates the model list (and picks Sonnet 4.6 vs 5 for the Claude_Sonnet tone); licenseType rides along and unlocks nothing by itself. |
M365_CACHE_FILE |
Override MSAL token cache location |
M365_SECRETS_FILE |
Override credentials file location |
CHROMIUM_PATH |
Path to Chromium binary for automated login |
The OpenAI usage block in every chat completion response now includes M365
extension fields with the per-conversation message quota — the closest
proxy we have to "context-window utilisation" since M365 hides token counts:
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0,
"x_m365_conversation_messages": 42,
"x_m365_conversation_max": 600,
"x_m365_conversation_pct": 7,
"x_m365_conversation_remaining": 558,
"x_m365_content_origin": "DeepLeo",
"x_m365_message_type": null,
"x_m365_turn_count": 3,
"x_m365_classifier_scores": {
"BotOffense": 1.27e-7,
"dea_violation": 2.81e-6
},
"x_m365_dea_score": 2.81e-6,
"x_m365_offense_score": 1.27e-7
}x_m365_dea_score is M365's own "disengaged-eligible answer" classifier
score — the closest signal to "am I about to get Disengaged?". Empirically:
clean tool calls sit at ~1 × 10⁻⁸, prose at ~1 × 10⁻⁶, jailbreak-shaped
prompts at ~1 × 10⁻³. Disengaged itself fires at some threshold > 2 × 10⁻³
that we haven't yet pinpointed. Clients can monitor this to back off before
tripping the filter.
Clients that ignore unknown extension fields keep working; curious users can read them. See docs/hypotheses.md §0 for the full findings dump and §2 for what we tried and didn't find.
All stored in ~/.config/opencode-m365/:
| File | Description |
|---|---|
secrets.json |
Login credentials (email, password, mfaSecret) |
msal-cache.json |
MSAL token cache (auto-managed) |
agent-id.json |
Cached Copilot Studio agent ID |
debug.log |
Debug log (when M365_DEBUG=1) |
pnpm install
pnpm build # Build all packages
pnpm run dev # Start standalone proxy on :4141
pnpm run test:unit # Run vitest unit tests (no auth/network)
pnpm run test:live # Run live integration tests against M365- M365 "disengages" on large tool payloads — heavy agent harnesses (e.g. opencode's ~15-tool prompt) get empty
Disengagedresponses. Keep the toolset lean (this is why pi works well). A heavy harness can be trimmed to fit, though opencode in particular needs that done in the proxy rather than through its own config, and its tool count has since dropped — see docs/m365-copilot-api.md. - Tool calling is emulated (prompt injection + a Copilot Studio agent), not native function calling — robust with the agent, unreliable without it
- The
think-deeper/*_Reasoningmodels take 10-30s per response - Hard quota of ~600 messages per conversation (mitigated by session reuse + delta sends)
claude-opusneeds a paid/premium seat and has its own small priority-access quota that resets at midnight UTC (weekly on Monday); exhaustion surfaces as a 429, not as a model answer- Streaming: tool-less responses stream incrementally (deltas forwarded as they arrive). Tool-calling turns are still buffered server-side — the raw text has to be parsed for tool-call fences before it can be emitted — so those arrive as a single chunk at the end (with an immediate HTTP 200 + heartbeats so the client never times out waiting)
MIT. Use at your own risk — this speaks to Microsoft's API with your own credentials, on your own account, and that's between you and your tenant's acceptable-use policy.