3.0 PREVIEW

Notice: GETSSH 3.0 PREVIEW is scheduled for release around October 10, 2026. Capabilities marked as "Coming soon" will be available for early access in this milestone.

MODELS & SPECS

Model Adaptation & BYOK Specification

3.0 PREVIEW

Deep protocol adaptation matrix, CoT isolation, and zero-store standards for mainstream LLMs.

This document serves as the authoritative reference for GETSSH users, developers, and security audit teams, systematically covering currently integrated providers, models under active internal evaluation, and those explicitly excluded from the roadmap — along with deep protocol adaptation specs, BYOK zero-retention guarantees, and streaming state machine design.


Table of Contents

  1. Adaptation Matrix Overview & Admission Criteria
  2. Core Principle: Zero Data Retention (BYOK Privacy Guarantee)
  3. Why Reject "Naive OpenAI Wrapper"?
  4. Unified Abstraction: Typed Block State Machine
  5. Currently Deep-Adapted Models
  6. Models Under Internal Evaluation
  7. Models Not Currently Supported
  8. Streaming Tool Call Assembly Matrix
  9. Implementation Checklist & Error Routing

1. Adaptation Matrix Overview & Admission Criteria

As a production-grade AI Agent terminal for infrastructure operations, GETSSH applies far stricter criteria to model admission than typical chat applications. The internal team continuously evaluates mainstream and emerging LLMs, making inclusion decisions based on the following criteria.

Admission Criteria

  • Reliable Function Calling / Tool Use: Multi-turn tool calls preserve full context; argument formats are stable and parseable; they don't silently regress across model versions.
  • Stable Streaming Output: SSE streams deliver in order without silent packet loss; recoverable on disconnection; terminal boundary events trigger correctly.
  • Agent-Grade Multi-Step Reasoning: Capable of maintaining task state across complex ops flows (e.g., analyze crash logs → draft recovery plan → human approval → execute) without context loss.
  • BYOK Privacy Compatibility: Supports direct API Key routing to official vendor endpoints with no forced server-side session retention and no mandatory vendor account prerequisite.
  • Protocol Stability: Mature API design with infrequent breaking changes; adapter maintenance cost remains controllable.

Provider Tiers

StatusDescription
✅ Deep AdapterFull custom adapter: CoT isolation, tool calling, BYOK, error routing
🔬 Internal TestingActive protocol evaluation; planned for a specific release
⏸ Not Currently SupportedAgent capability or protocol maturity does not yet meet production-grade admission bar
🌐 Generic FallbackAny /v1/chat/completions-compatible custom endpoint can be connected, without targeted optimization

2. Core Principle: Zero Data Retention (BYOK Privacy Guarantee)

GETSSH is engineered as a local-first terminal and infrastructure management environment. Terminal operations inherently involve sensitive production topologies, host credentials, and system logs.

BYOK Architecture Rules

  • Bring Your Own Key, Direct Official Endpoints: User API keys are securely stored in OS-level hardware vaults (macOS Keychain / Secure Enclave, Windows DPAPI, Linux Secret Service) and injected directly to vendor endpoints via encrypted TLS.
  • Enforced Server Storage Disablement (store: false):
    • OpenAI Responses API defaults to persisting sessions (store: true). GETSSH explicitly enforces store: false across all outbound requests.
    • Google Gemini Interactions API also dispatches store: false explicitly.
    • User host inventories, SSH session logs, and diagnostic outputs are never retained on third-party servers.
  • Zero Intermediary Proxies: Unless custom proxies are explicitly configured by the user, all traffic is peer-to-cloud with no proprietary gateway telemetry.

3. Why Reject "Naive OpenAI Wrapper"?

Many clients attempt to wrap all models under OpenAI's Chat Completions interface. In production-grade terminal Agent scenarios, this causes severe functional degradation and execution lockups:

ScenarioNaive Wrapper FailureGETSSH Deep Custom Adapter
Thinking ChainsCoT mixed into delta.content, flooding the terminalIsolates reasoning_content / thought, rendering folded ambient status
Tool ResultsMisusing id instead of call_id or losing signaturesExact pairing: call_id (OpenAI), signature (Claude), thought_signature (Gemini)
Locked HyperparametersPassing temperature to Claude / Kimi returns HTTP 400Automatically strips unsupported parameters for 100% request success
Streaming FragmentsAssuming tool args arrive as complete JSONBuffers fragments with keyed maps, parsing only on terminal boundary events
Quota vs Rate LimitsInfinite retries on pseudo-429s (unpaid balance)Discriminates 1113 (GLM), 402 (DeepSeek), billing_error (Anthropic) immediately

4. Unified Abstraction: Typed Block State Machine

To bridge distinct vendor protocols without data loss, GETSSH models all turns internally as Ordered Typed Blocks:

export type Block =
  | { kind: 'text'; text: string }
  | { kind: 'thought'; text?: string; opaque: string } // opaque: verbatim signature or encrypted payload
  | { kind: 'tool_call'; callId: string; name: string; args: unknown; raw: string }
  | { kind: 'tool_result'; callId: string; content: Block[]; isError?: boolean }
  | { kind: 'image'; mime: string; data: string }
  | { kind: 'server_tool'; toolType: string; payload: unknown; opaque?: string };

export interface Turn {
  role: 'user' | 'assistant';
  blocks: Block[];
}
  • Verbatim opaque Signatures: Cryptographic signatures (encrypted_content, signature, thought_signature) are transmitted back untampered to maintain reasoning integrity across tool turns.
  • Original Parameter Preservation: Preserves raw argument strings to handle out-of-order parallel tool streams safely.

5. Currently Deep-Adapted Models

OpenAI (Responses API)

  • Primary Endpoint: POST https://api.openai.com/v1/responses.
  • Key Adaptations:
    • Chat Completions cannot combine reasoning with tool calls. GETSSH routes all tool-assisted operations via Responses API.
    • Tool execution returns must correlate using call_id (not id).
    • Reasoning effort levels: low, medium, high, xhigh, max.
    • Always enforces store: false for BYOK privacy compliance.

Anthropic (Claude)

  • Primary Endpoint: POST https://api.anthropic.com/v1/messages, required header anthropic-version: 2023-06-01.
  • Key Adaptations:
    • Strict Hyperparameter Guard: Non-default temperature, top_p, or top_k trigger an immediate HTTP 400 on modern models. GETSSH strips these by default.
    • Adaptive Thinking: Thinking blocks include an encrypted signature. In subsequent tool turns, both thinking and redacted_thinking blocks must be preserved in exact original order.
    • Tools declared under input_schema (not parameters).

Google Gemini (Interactions API)

  • Primary Endpoint: POST https://generativelanguage.googleapis.com/v1beta/interactions (default since June 2026, replacing generateContent).
  • Key Adaptations:
    • Step-based model output: model_output, function_call, thought.
    • Tool execution state signaled via status: "requires_action".
    • In stateless BYOK mode, thought step and thought_signature must be re-transmitted verbatim in subsequent requests.

DeepSeek (Single-Model Hybrid Reasoning)

  • Primary Endpoint: https://api.deepseek.com/chat/completions.
  • Key Adaptations:
    • Hybrid Architecture: Single unified model dynamically controlled via thinking: {type: "enabled"}.
    • Mandatory CoT Historical Replay: If tools are present, reasoning_content from all prior turns must be transmitted back, even for turns that did not invoke tools.
    • Concurrency Semaphore & Keep-Alive: Client-side concurrency control and automatic filtering of : keep-alive SSE comment lines.

Zhipu GLM (Mandatory Thinking)

  • Primary Endpoint: https://open.bigmodel.cn/api/paas/v4/chat/completions.
  • Key Adaptations:
    • Modern Bearer authentication (Authorization: Bearer <key>) replacing legacy JWT.
    • Mandatory Thinking: Permanently enabled; cross-turn preservation via thinking.clear_thinking: false.
    • temperature range restricted to [0.0, 1.0]. Deterministic outputs use do_sample: false.
    • tool_choice supports "auto" only; explicit enforcement is managed client-side.

Moonshot Kimi (1M Context)

  • Primary Endpoint: https://api.moonshot.cn/v1/chat/completions.
  • Key Adaptations:
    • Locked Sampling Parameters: Passing temperature, top_p, n, presence_penalty, or frequency_penalty returns HTTP 400. All outbound requests are automatically sanitized.
    • Primary limit field: max_completion_tokens (not max_tokens).
    • $web_search builtin protocol: verbatim argument echo without local execution.
    • Flat usage.cached_tokens field for accurate ops cost accounting.

Tongyi Qwen (Aliyun DashScope Native Mode)

  • Primary Endpoint: POST https://dashscope.aliyuncs.com/api/v1/services/aigc/text-generation/generation (streaming header: X-DashScope-SSE: enable).
  • Key Adaptations:
    • Native DashScope Mode Enforced: OpenAI-compatible mode fails to return web search attribution (search_info). GETSSH uses the native protocol to surface citations in server_tool blocks.
    • Mandatory Incremental Output: Enforces parameters.incremental_output: true and result_format: "message" to prevent cumulative text repainting.
    • Tool Choice Auto-Downgrade: In enable_thinking: true mode, forced tool calls are rejected by the provider; GETSSH automatically degrades tool_choice to "auto".
    • Custom SSE Parser: Handles DashScope's compact id:1 and event:result lines while filtering :HTTP_STATUS/200 comment frames.

MiniMax (Dual-Protocol & False-Success Shield)

  • Primary Endpoint: POST https://api.minimax.io/anthropic/v1/messages (fallback: /v1/chat/completions).
  • Key Adaptations:
    • Anthropic Protocol by Default: Standard HTTP codes with native structured thinking blocks, sidestepping the base_resp pseudo-200 pitfall.
    • OpenAI Mode Reasoning Split: When routed via /v1/chat/completions, injects reasoning_split: true to prevent CoT leaking into content as raw <think> tags.
    • Thinking Toggle Dispatch: Permits thinking: { type: "disabled" } on MiniMax-M3 while guarding M2.x models against invalid disable requests.
    • base_resp Circuit Breaker: Deep inspection captures 1004 (auth failure) and 1005 (insufficient balance) immediately, halting retry storms.

6. Models Under Internal Evaluation

🔬 xAI Grok

Current Status: Internal protocol testing in progress. Expected to ship with GETSSH 3.2 Fusion Beta Preview.

Grok models (including Grok 4, Grok 4.6 and frontier releases by xAI) demonstrate cutting-edge real-time awareness and deep reasoning capabilities, and the open API shows a comparatively mature approach to tool calling pipeline design. The internal team is currently running systematic evaluation across the following dimensions:

  • Multi-turn tool call argument integrity and call_id correlation consistency.
  • Extended Thinking streaming segmentation behavior and cross-turn relay protocol.
  • BYOK privacy compatibility: identifying any server-side forced session retention paths and equivalent store: false controls.
  • Multi-step reasoning stability in GETSSH production ops scenarios (cross-host command orchestration, log analysis, rolling restart workflows).

Full protocol adapter documentation will ship alongside the 3.2 Fusion Beta Preview release.


7. Models Not Currently Supported

The internal team continuously tracks mainstream and emerging LLMs, conducting systematic evaluations when conditions are right. The following providers or products currently do not meet GETSSH's production-grade admission bar across Agent capability maturity, multi-turn tool call stability, or protocol standardization. They are not on the official supported adapter list at this time.

This is not a negative assessment of their underlying technology. The admission list will evolve dynamically as the LLM ecosystem advances.

Mistral / Mistral AI

Mistral's Agents API is at a relatively early stage. Multi-step tool call chain stability in complex ops workflows requires further validation. The language models have notable limitations in handling Chinese system logs, mixed-language error outputs, and code-switching command sets — which are core GETSSH use cases. The team is monitoring; evaluation will continue as the API matures.

Baidu ERNIE (Wenxin Yiyan)

The internal team ran a dedicated evaluation on ERNIE 5.0. The conclusion: even at the latest version, its performance in the complex Agent scenarios defined by GETSSH does not meet the bar — multi-turn tool call context consistency and mid-chain self-recovery capability both have significant gaps that prevent production-grade ops Agent admission. The team will continue tracking future releases, but will not ship an official adapter until there is a meaningful capability breakthrough.

ByteDance Doubao

Internal evaluation found that Doubao models exhibit pronounced hallucination behavior in ops Agent scenarios, along with a tendency toward biased outputs on infrastructure-related instructions and contextual signals — risks that are unacceptable in production environments. The team will continue monitoring its technical evolution, but will not release an official adapter until hallucination rates and output reliability reach an acceptable threshold.

iFlytek Spark

iFlytek Spark's product positioning is as a general-purpose conversational assistant and vertical industry language application — its architecture is not designed for Agent automation task chains. In evaluation, its performance on core Agent capabilities — multi-step tool orchestration, execution state persistence, and error recovery — trails models built specifically for Agent use by a fundamental architectural margin. This is a directional gap that version iteration alone cannot close. Not in the evaluation queue.

Meta Llama (Self-Hosted)

Llama models deployed locally via Ollama, vLLM, LM Studio, or similar runtimes can be connected directly through GETSSH's built-in generic OpenAI-compatible adapter — no specialized setup required. The official Llama API (llama.ai) has unstable tool calling specifications at this time; no dedicated adapter layer is planned.

Other Providers (Cohere, AI21 Labs, etc.)

These international providers have made limited investments in the AI Agent / Function Calling domain. Their specialized capabilities in complex infrastructure operations scenarios trail the mainstream tier significantly. Not currently in the evaluation priority queue.


On Generic Fallback: Any OpenAI /v1/chat/completions-compatible local endpoint (Ollama, vLLM, LM Studio, self-hosted gateways, etc.) can be connected to GETSSH as a custom endpoint via the generic compatibility adapter. This path is functional, but does not benefit from targeted CoT isolation, parameter safety filtering, or intelligent error routing.


8. Streaming Tool Call Assembly Matrix

ProviderChunk KeyDelta FieldStop EventFinal Parse Point
OpenAI Responsesitem_iddelta*.doneOverwrite with .done full object
OpenAI CC (Legacy)tool_calls[].indexfunction.argumentsfinish_reason: "tool_calls"Parse upon stream termination
Anthropic Messagesblock indexdelta.partial_jsoncontent_block_stopTrigger parse at content_block_stop
Gemini Interactionsstep indexdelta.argumentsstep.stopTrigger parse at step.stop
DeepSeektool_calls[].indexfunction.argumentsfinish_reason: "tool_calls"Parse upon stream termination
Zhipu GLMtool_calls[].indexfunction.argumentsfinish_reason: "tool_calls"Parse upon stream termination
Moonshot Kimitool_calls[].indexfunction.argumentsfinish_reason: "tool_calls"First content chunk concludes thinking
Tongyi Qwentool_calls[].indexfunction.argumentschoices[0].finish_reasonCustom compact SSE parser, first content ends thinking
MiniMaxblock index (Anthropic) / item_iddeltamessage_stop / finish_reasonTraps pseudo-200 base_resp, splits reasoning

9. Implementation Checklist & Error Routing

P0 Critical Guardrails

  1. Thinking Integrity: Never strip or edit thinking signatures in transit.
  2. Sanitize Sampling Parameters: Strip unsupported parameters for Claude and Kimi.
  3. OpenAI Route Switching: Always dispatch via Responses API when reasoning and tools co-exist.
  4. Enforce BYOK Zero-Store: Force store: false across all outbound calls.

Intelligent Error Discrimination

  • Intercept Non-Retryable Account Errors:
    • GLM code 1113 (unpaid account pseudo-429): surface a top-up alert immediately, block retry;
    • DeepSeek 402 Insufficient Balance: hard circuit break;
    • Kimi exceeded_current_quota_error: terminate session;
    • MiniMax 1005: immediate circuit open, halt retry storm.
  • Adaptive Exponential Backoff: Applied solely to bona fide rate limits (rate_limit_error, GLM 1302/1305, DeepSeek 429) to maintain smooth ops diagnostic pipelines.