Skip to content

Compose Semantics Inspector ​

The Compose Semantics Inspector is an official JetWhale plugin that reads the semantics tree of your running app — every node, its labels, its bounds, and the actions it exposes — and lets you both browse it in the host and hand it to an AI agent over MCP. On Android and desktop that is the Compose tree, with the Android Views around it; on iOS it is the accessibility tree, which carries UIKit, SwiftUI and Compose content alike.

  • ⚡ ~14 ms per capture end-to-end, against ~2.7 s for android layout — see Why not the CLI?
  • 🌲 Live tree of the app's nodes, with search and an interactive only filter
  • 🎯 Per-node detail: role, text, contentDescription, testTag, state, bounds in root and on screen — in pixels on Android and desktop, in points on iOS
  • 👆 Run a node's own semantics action — click, long click, set text, scroll, focus, dismiss — from the host or from an AI agent
  • 🪟 On Android and desktop, dialogs and popups appear as their own roots, because that is what they are in Compose; on iOS a sheet or alert stays inside the window that showed it
  • 🤝 On Android, the Android Views around and inside the composition are in the same tree, and a View's attributes can be read and edited live — see Android View support; on iOS, UIKit and SwiftUI are read alongside Compose — see iOS support
  • 🖐 Whether the user could operate a node right now — enabled, and reached by a gesture it accepts — worked out from the capture rather than by tapping — see Can the user operate it?
  • 🤖 Six MCP tools so an agent can see the screen structurally instead of guessing at pixels

What the tree contains ​

The tree is the semantics tree: the same tree an accessibility service sees, and the one that says what is actually clickable. A Box that only lays out pixels does not appear on its own; a Button does, carrying its label and its OnClick action. On Android and desktop it is read from Compose, and each root is a window in screen pixels with dialogs and popups as roots of their own; on iOS it is read through the accessibility protocol, each root is a UIWindow in points, and a dialog stays inside the window that showed it. Every node in the MCP JSON names its unit.

Two views of it are available, switchable in the host and per MCP call:

What you get
Merged (default)A Button's label is folded into the clickable node — one node per control. This is what accessibility services and performClick see, and usually what you want.
UnmergedEvery semantics node stays separate, closer to how the UI is written. Useful when you need to see exactly which composable contributed which property.

On Android the tree also carries the Views around and inside the composition; on iOS it is read through the accessibility protocol and carries UIKit and SwiftUI content alongside Compose. See Android View support and iOS support.

Android View support ​

Very little Compose is only Compose. A screen usually sits in an Activity's or a Fragment's layout, and an AndroidView { } puts a View back inside the composition. On Android the capture follows both crossings, so a root is one window, not one composition:

MainActivity                              ← the window's decor view
└─ … the layout around the ComposeView …
   └─ AndroidComposeView                   ← where the composition starts
      └─ Column                            ← Compose semantics from here down
         ├─ Button · Send
         └─ LinearLayout                   ← an AndroidView { }, and its subtree
            ├─ TextView · @id/status
            └─ Button · @id/submit

A View node is its own node type on the wire — "type": "view", against "compose" for a semantics node — and it fills the same fields a Compose node does: text, contentDescription, bounds, actions, and the enabled/clickable/editable/scrollable flags. So nothing that only reads the tree has to special-case it. What is particular to it:

View node
idnegative, assigned by the agent and valid while the view is alive — Compose's semantics ids are non-negative, so the two can never collide
viewClassthe view's class, e.g. android.widget.Button
resourceIdthe entry name of its android:id, e.g. submit for @id/submit — a View has no testTag, and this is what plays that role

A Compose node carries role, testTag and stateDescription, which a View has no counterpart for; a View node carries viewClass and resourceId, which a Compose node has no counterpart for.

The two node types are told apart on the wire by "type": "compose" / "type": "view", so host and agent have to be built against the same protocol version — a mismatch fails to decode rather than degrading.

performNodeAction works on a View node too, running the view's own API rather than a synthesised tap: Click → performClick(), LongClick → performLongClick(), SetText / InsertText on an EditText, ImeAction → onEditorAction, ScrollBy → scrollBy, ScrollToIndex → RecyclerView.scrollToPosition / ListView.setSelection, BringIntoView → requestRectangleOnScreen, RequestFocus → requestFocus(). Dismiss, Expand and Collapse have no View counterpart and come back performed: false saying so. As always, only what a node lists in actions can be invoked — except BringIntoView, which every node takes.

Editing View attributes ​

A View node also exposes its platform attributes — the properties a Layout Inspector shows — and most of them can be changed live, so you can try a padding, a color or a GONE without a rebuild. Select a View node in the host and the attributes appear under its semantics, grouped:

GroupAttributes
Statevisibility (VISIBLE/INVISIBLE/GONE), enabled, selected, activated, clickable, focusable, and focused read-only
Layoutlayout.width / layout.height (MATCH_PARENT, WRAP_CONTENT or a pixel length), padding.*, margin.* (when the parent hands out margins), minWidth / minHeight, and bounds read-only
Appearancealpha, backgroundColor (when the background is a flat color; otherwise background names the drawable, read-only), elevation, translationX / translationY, rotation, scaleX / scaleY
Texton a TextView: text, hint, textSize, textColor, maxLines
Infoid (@id/name) and class, both read-only

Each attribute has one type, and the type is what decides the editor — a visibility is always a dropdown of its options, a padding.left always a pixel field. layout.width and layout.height take either a constant or a length, so they have a type of their own that says both: a dropdown of MATCH_PARENT, WRAP_CONTENT and Fixed, with a pixel field that Fixed enables. The editor looks the same whichever the size currently is, so a WRAP_CONTENT can be given a width and a width can be put back to WRAP_CONTENT.

Three things to know:

  • It is a curated list, not reflection. Every attribute is named explicitly. Reflection over a view's getters is what makes the equivalent elsewhere fragile, and Android's non-SDK interface restrictions block most of what it would reach anyway. The list says exactly what the agent will touch.
  • An edit is temporary. The app owns the property: a relayout, a rebind, or the app writing it itself takes the value back. This is a way to see a change, not to make one.
  • Compose nodes have none. A semantics node is a projection of composition state, so writing to it would be undone by the next recomposition — which is why Android Studio's Layout Inspector edits views and not Compose either. Asking for the attributes of a Compose node answers a message saying so rather than failing.

Attributes are fetched per node when you select one, never as part of a capture: a tree of two hundred nodes must not carry thirty attributes each.

Two MCP tools do the same from an agent — getViewAttributes and setViewAttribute.

Two limits are worth knowing:

  • A window with no Compose in it is not captured. The composition is what announces a window to the probe, so a plain AlertDialog built from views does not appear. This plugin inspects Compose apps; it is not a general View inspector.
  • A merged capture can fold an AndroidView away. The embedded views hang off the semantics node the AndroidView { } creates; when an ancestor merges its descendants (a Button, a mergeDescendants = true modifier), that node is folded into the ancestor in the merged tree and the views under it go with it. Capture unmerged (merged: false) to see them.

Other platforms are unaffected: desktop reads a composition through its SemanticsOwner, and its roots stay one-per-composition.

iOS support ​

On iOS the capture does not read Compose through a SemanticsOwner — Compose Multiplatform hands none out for the scene an app shows — but through the accessibility protocol, which every toolkit on the platform publishes into: UIKit views are their own accessibility objects, SwiftUI lists its nodes through its hosting view, and Compose Multiplatform lists its semantics tree from the view it draws into. One walk per window covers all three, whether the app is Compose with a SwiftUI bar above it, SwiftUI with a ComposeUIViewController inside, or plain UIKit:

UIWindow                                  ← the root; one per app, dialogs and sheets included
└─ _UIHostingView                         ← SwiftUI
   ├─ AccessibilityNode · swiftui-button  ← a SwiftUI Button
   ├─ UITextField · swiftui-name          ← a SwiftUI TextField is a real UITextField
   └─ ComposeContainerView                ← where Compose starts
      └─ AccessibilityRoot
         └─ AccessibilityElement · increment-button   ← a Compose Button, testTag and all

Every node on iOS is a "type": "apple" node — there are no "compose" nodes, because the Compose content arrives through the same protocol as everything else. It fills the same fields a Compose node does, so a consumer that only reads the tree needs nothing new. What is particular to it:

apple node
idnegative, assigned by the agent and valid while the object is alive
classNamethe Objective-C class: UITextField, SwiftUI.AccessibilityNode, Compose's AccessibilityElement
accessibilityIdentifierwhere SwiftUI's .accessibilityIdentifier(_:) and a Compose Modifier.testTag both land — findNodes(testTag:) matches it
accessibilityValuethe value as the toolkit reports it: a switch's "1", a slider's "50%"
traitsthe set UIAccessibilityTraits, by name: Button, Selected, NotEnabled, ToggleButton…

Coordinates are points, the unit every iOS tool takes — idb ui tap X Y from idb takes them as they are — and the root's density is 1. Every node in the MCP JSON says which unit its bounds are in — "unit": "pt" here, "px" elsewhere — so a flat findNodes result needs no root to read them. merged has no effect: the accessibility tree is the merged one, and it is the only one there is.

A view marked accessibilityViewIsModal — a presented sheet, an alert — hides its siblings the way it hides them from VoiceOver: what is behind the modal is invisible, so a default capture leaves it out and includeInvisible brings it back, marked as such.

What the accessibility protocol does not carry, the capture cannot report: a Compose role and stateDescription are folded into label and traits, and a node scrolled out of its container reports its laid-out frame, clipped to the window rather than to the container. A Compose scrollable does read scrollable: its element conforms to UIKit's UIFocusItemScrollableContainer exactly when the node scrolls, which also gives ScrollBy a real distance there.

performNodeAction runs the accessibility protocol's counterpart, or the view's own API when the node is a view that has one:

ActionUIKit viewSwiftUI nodeCompose element
ClickaccessibilityActivate(), else the control's touch-up actionsaccessibilityActivate() — runs the Button's closure, flips a ToggleaccessibilityActivate() — runs onClick, toggles a Checkbox
SetText / InsertTexton a UITextField / UITextView; InsertText focuses the field first, as a keystroke needssame: the TextField is a UITextField underneathnot available — the element's value is read-only
ImeActionthe field's delegate textFieldShouldReturn:same — where onSubmit livesnot available
ScrollByUIScrollView.setContentOffset, by the distance askeda ScrollView is a UIScrollView underneath; a bare node gets accessibilityScroll, by direction: one pageby the distance asked, through the UIFocusItemScrollableContainer the element conforms to
ScrollToIndexUITableView / UICollectionView, the index counted across sectionsa List is a UICollectionViewnot available
BringIntoViewevery UIScrollView above the viewa bare node only reports whether it is already in viewsame
RequestFocusbecomeFirstResponder()on the backing UITextFieldnot available
DismissaccessibilityPerformEscape(), tried on any nodesamesame
Expand / Collapsea custom action of that name, when it has a handler block; a target/selector custom action is not invokedsamesame
LongClicknot available

Not covered on iOS: LongClick; SetText / InsertText / ImeAction / RequestFocus / ScrollToIndex on a Compose element; custom actions built with a target and selector rather than a handler block; and text entry into a secure field's contents, which are never captured either: a password field stays isEditable with no editableText, on Android as on iOS.

Setup ​

Install the host plugin ​

The Compose Semantics Inspector is in the host's official catalog: open Settings → Plugins → Add Plugins → Official Plugins and install it with one click — no coordinates needed. See Host Settings → Plugins for the other install routes.

Add the agent to your app ​

Add the agent to the app being debugged — one artifact, carrying both the plugin and the probes (see Platform support for which targets have a probe):

kotlin
dependencies {
    implementation("com.kitakkun.jetwhale:jetwhale-agent-runtime:<version>")
    implementation("com.kitakkun.jetwhale:jetwhale-compose-semantics-inspector-agent:<version>")
}

Register the plugin with the agent runtime, and install a probe so it has roots to read.

Registering the plugin ​

kotlin
import com.kitakkun.jetwhale.agent.runtime.startJetWhale
import com.kitakkun.jetwhale.plugins.semantics.agent.JetWhaleSemanticsAgentPlugin

startJetWhale {
    connection {
        endpoints {
            ws("localhost", 5080)
        }
    }
    plugins {
        register(JetWhaleSemanticsAgentPlugin())
    }
}

This part is common code — it compiles on every Compose Multiplatform target. Without a probe the plugin still answers, reporting an empty tree and a warning saying so, which the host shows.

Installing a probe ​

There are two ways in, and they can be combined — registrations are reference counted per window, so neither can pull a window out from under the other.

kotlin
import com.kitakkun.jetwhale.plugins.semantics.agent.installJetWhaleSemanticsProbe

class MyApplication : Application() {
    override fun onCreate() {
        super.onCreate()
        installJetWhaleSemanticsProbe(this)
        startJetWhale { /* … */ }
    }
}

Installed from onCreate(), before any activity exists, it hooks the callback Compose fires when it creates the view backing a composition — so it sees every window holding one, including the separate one a Dialog or a Popup opens. No screen has to change.

Installed later it also scans the resumed activity's window, so roots created before the call are not lost — but roots in windows that were already open and are never re-resumed can be.

The install is process-wide and idempotent, and closing the returned handle restores whatever callback was there before — so it composes with the Compose test framework rather than displacing it.

From inside the composition ​

kotlin
import com.kitakkun.jetwhale.plugins.semantics.agent.JetWhaleSemanticsProbe

setContent {
    JetWhaleSemanticsProbe()
    App()
}

Use this when the Application layer is not yours to touch, or when only one screen should be readable. It registers its own root for as long as it stays composed — on Android, the window that root lives in. A Dialog or Popup is a root of its own, so add a call inside those too, or install the Application-level probe, which finds them all.

On iOS ​

kotlin
import com.kitakkun.jetwhale.plugins.semantics.agent.installJetWhaleSemanticsProbe

installJetWhaleSemanticsProbe()

Call it once at startup, on the main thread — before or after startJetWhale. It registers every window the app has and follows UIWindowDidBecomeVisible / UIWindowDidBecomeHidden for the ones that open later; there is nothing per screen to add, and no composition to put a call inside. A pure Swift app cannot call it yet: the agent's start API has no Swift surface, so the call has to sit in Kotlin the app links, as in the demo's cmpAppViewController().

Debug builds only

The probe makes your app's UI structure readable, and the actions below make it drivable, over the JetWhale connection. Wire both up in debug builds only, exactly as you would the rest of JetWhale.

Using the host UI ​

Open the Compose Semantics Inspector in the host, select your app's session, and press Refresh.

  • Auto re-captures once a second. It is off by default: a capture reads the app's semantics on its main thread, so leaving it on makes the app do that work forever.
  • Interactive only keeps the nodes that expose an action, are editable, or scroll — plus their ancestors, so the structure stays readable.
  • Include invisible adds nodes that are not laid out or are fully clipped away.
  • The search box matches text, contentDescription, testTag, role and id.

Select a node to see its full semantics on the right, along with a button for every action it actually exposes. There is also a Copy adb shell input tap button — Copy idb ui tap for an iOS node — for the times you do want to drive the app through the input system. Select an Android View node and its editable attributes appear below that — see Editing View attributes.

Highlighting on the device ​

Turn on Highlight and the app draws a translucent box over the node you have selected; hovering a row shows that node instead for as long as the pointer is on it, and the selection comes back when it leaves. It works for Compose nodes and Android View nodes alike — both report their bounds in the same window coordinates — and a node in a dialog is highlighted in the dialog's own window.

Four things to know:

  • It is off by default, on purpose. The box is drawn into the app itself, so anything that takes a screenshot of the device while it is up captures the box too — a screencap, the Android Device plugin, a QA run. Turn it on to find something, turn it off before you capture.
  • It never appears in the captured tree. The box is a window overlay (View.getOverlay()), drawn after the root view's children but not one of them, so the tree you are reading is not changed by reading it.
  • It follows the node, or goes away. The box is put back where the node is as the window redraws, so scrolling the app, a relayout, or a rotation the activity handles itself does not leave it behind; a node that can no longer be found takes the box down rather than stranding it, and a node that is only out of view for now — scrolled out of a lazy list, say — gets its box back when it comes back while still selected. A box on screen is always in the right place.
  • It clears itself. The app drops a highlight it has not heard about for 30 seconds, so a host that crashes or is killed cannot leave a box on the app's screen for the rest of the session. Closing the inspector, disabling the plugin and disconnecting all clear it straight away.

There is deliberately no MCP tool for this. An agent reads a node's bounds from the tree already, and a highlight would only put a box into the screenshots it takes.

Can the user operate it? ​

The question an agent has before it acts is one word: operable. A node is operable when it offers something to do (an action, editable content, or scrolling), is enabled, and a gesture it accepts actually reaches it. A node that offers something to do but fails one of those is marked "operable": false, with enabled, hittable and obscuredBy saying which; the tree view tags the row not operable. findNodes(operableOnly: true) keeps only the nodes that pass.

A label is never marked either way: there is nothing on it to operate.

Can a finger reach it? ​

The reachability half of that answer is its own flag. Every captured node says whether a gesture aimed at it would actually arrive. A node that accepts touch input but cannot receive one is marked "hittable": false, with obscuredBy naming what takes the touch instead.

It is worked out from the capture alone — no tap sent, nothing to wait for — by walking the windows from the top down and, within a window, the last-drawn node first, the way the platform dispatches a touch. That catches what is worth catching:

The button is…Reported as
under a dialog or a popuphittable: false, obscuredBy the node on top
behind a modal window, anywhere on screenhittable: false, obscuredBy that window's root
scrolled out of its containerhittable: false, no obscuredBy — no area to aim at
under a later sibling that takes toucheshittable: false, obscuredBy that sibling
a clickable row whose center is its own buttonhittable: false, obscuredBy that button — the tap is consumed there
a list whose center is one of its own rowshittable: true — a drag starting on the row still scrolls the list

The last two differ because the gesture does: a tap stops at the deepest node that takes it, while a drag is seen by every scrollable on the way down. So obscuredBy on a scrollable only ever names something outside it — a sibling drawn on top, or another window. A node that accepts both — a scrollable that is also clickable — reads hittable: true when either gesture gets through.

Two things a capture cannot see, and both make it optimistic — a node can read as reachable that a finger would not reach:

  • an overlay that consumes touches without exposing any semantics (a bare pointerInput, an OnTouchListener),
  • a gesture an ancestor swallows before the node sees it.

On the Compose side the order this relies on is guaranteed — semantics children come from the layout's z-sorted children, so Modifier.zIndex is accounted for. An Android ViewGroup is read in child order, which is paint order until a view is raised by elevation or translationZ; a raised sibling can be missed as an obstruction.

Both are rarer than they sound. Sweeping the demo app point by point — comparing what the tree predicts against what a real touch does — the tree was right at every interactive node; the only places the two parted company were empty ones, where Material's Surface takes a touch that no node was going to get anyway. Where it matters, treat hittable: false as reliable and hittable: true as "nothing in the tree is in the way".

Note that none of this constrains performNodeAction, which invokes the node's own action and never goes near the input system. Hittability is about whether a person could tap it — a UI check, not a precondition for driving the app.

Why not just tap it and see? ​

Because a tap is not a question. It fires the action, changes the screen, and still does not say what it hit — finding that out means dumping the hierarchy afterwards and undoing whatever happened.

Measured on the same emulator and screen as the capture benchmarks — the demo app's Compose nodes screen, 27 nodes of which 8 are interactive:

answersmedian
adb shell input tapnothing — it only fires329 ms
adb shell input tap + uiautomator dumpone point, having changed the screen3,621 ms
nodeAtone point, from the tree30 ms
getNodeTreeevery node's reachability at once30 ms

One capture carries the whole screen's answer, so per interactive node it is about 4 ms — against 329 ms per point for a tap that also has to be undone.

The ratio is the smaller half of it. What the numbers buy is frequency: at 30 ms an agent can ask after every action, which at three seconds it cannot. And the answer is a node with its testTag, role and id — something to act on next — rather than a rectangle.

Treat all of these as indicative. They come from an emulator, which is slower than a device, and a bigger screen means more nodes.

MCP tools ​

The plugin contributes six tools to the host's MCP server. As with every plugin tool, JetWhale injects the sessionId parameter and routes the call to the right session.

com.kitakkun.jetwhale.semantics.findNodes ​

The one to reach for first. Captures the tree and returns the matching nodes as a flat list, each carrying the rootId/id pair that addresses it, its bounds on screen, and a ready-made tap point. Both are in the node's unit: pixels on Android and desktop, points on iOS.

Criteria (text, contentDescription, testTag, resourceId, role) are combined with AND and match case-insensitively by substring unless exact is set — resourceId is the exception, always compared whole, because a resource id is an identifier rather than a label. With no criteria at all it lists everything interactive on screen — a good way to answer "what can I do here?". Add operableOnly: true to narrow that to what the user could operate right now — see Can the user operate it?.

An Android View node is marked with "kind": "View" and carries its viewClass and resourceId — see Android View support.

com.kitakkun.jetwhale.semantics.getNodeTree ​

The whole tree, structure included. Takes merged, includeInvisible, maxDepth, interactiveOnly, rootId and format (see Compact text output). Use it when the layout itself is the question; use findNodes when you are looking for one element.

Compact text output ​

getNodeTree and findNodes return JSON unless the call passes format: "text". Then the same nodes come back as an outline, one line per node, indented by depth, for an agent that reads the tree rather than parses it:

text
root compose-root-1f2e "MainActivity" unit=px density=2.0
- node #1 tap=540,1200
  - node #2 tap=540,100
    - Button #3 desc="Navigate up" [clickable] actions=Click tap=80,120
    - node #4 "Settings" tap=390,120
  - node #5 tag=settings-list [scrollable] actions=ScrollBy,ScrollToIndex tap=540,1200
    - node #100 [clickable] actions=Click tap=540,370
      - node #200 "Setting item number 0" tap=374,345
      - node #300 "Description for item 0" tap=374,390
      - Switch #400 [clickable] actions=Click tap=970,370
    …

Each line starts with the node's role (a Compose node), its class (an Android View or an iOS node), or node when it has neither, then #id — with the root line's rootId, the pair performNodeAction takes. What follows is only what is present: the text in quotes, desc=, input= (a text field's content), toggle=, tag= / resId= / axId=, the surprising side of each flag in brackets ([clickable], [disabled], [unhittable], …), actions= (the names performNodeAction accepts, as in the JSON), and the tap point in the root's unit. Text is quoted as a JSON string, so a quote or a line break in it cannot break the line. Bounds are left out; ask for JSON when a layout question needs them. A findNodes line adds root=<rootId> after the id, since a flat list has no root lines.

On a 67-node settings screen the outline is 3.8 KB where the JSON is 10 KB.

com.kitakkun.jetwhale.semantics.nodeAt ​

Which node a tap at a screen coordinate would be dispatched to, or null when nothing there takes touch input. The question you have once you have picked a point from a screenshot rather than from a node's own bounds — see Can a finger reach it?.

com.kitakkun.jetwhale.semantics.performNodeAction ​

Invokes a node's own semantics action: Click, LongClick, SetText, InsertText, ImeAction, ScrollBy, ScrollToIndex, RequestFocus, Dismiss, Expand, Collapse — plus BringIntoView, which is not the node's own but works on any node. On an Android View node it runs the view's own equivalent — see Android View support; on iOS it runs the accessibility protocol's — see iOS support.

This runs the action the node itself declared, so it needs no coordinates and cannot land on whatever moved into that spot in the meantime — prefer it over adb shell input tap. rootId is optional: without it the node is looked up in the most recent capture.

It is also the more reliable route, not just the tidier one: on the emulator used for the benchmarks, adb shell input swipe did not scroll a LazyColumn at all, while ScrollBy moved it by exactly the requested distance on the first try.

A typical agent loop:

findNodes(testTag: "login-button")     → { "nodes": [{ "rootId": "compose-root-1f2e", "id": 42, … }] }
performNodeAction(nodeId: 42, action: "Click")
findNodes()                            → the new screen's interactive nodes

Getting a node on screen ​

A node that exists but sits outside the viewport is a poor target for a screenshot or a finger. BringIntoView scrolls it in the way accessibility's "show on screen" does: every scrollable ancestor is scrolled by the least amount that reveals the whole node, innermost first, and on Android the window's Views around the composition are scrolled too — so it crosses a LazyColumn inside a ScrollView, or a RecyclerView inside an AndroidView { }, in one call. No gesture, no coordinates, and a node that is already fully visible reports performed: true with a note saying nothing moved.

findNodes(text: "Terms of service")    → { "nodes": [{ "id": 87, "isVisible": false, … }] }
performNodeAction(nodeId: 87, action: "BringIntoView")
getNodeTree()                          → node 87 now has on-screen bounds

BringIntoView needs the node to exist, and a lazy container only composes the items near its viewport — an item far down a LazyColumn has no node yet. ScrollToIndex is the step before: invoke it on the container with the item's index, then capture again and the item is there. LazyColumn, LazyRow, the lazy grids and Pager expose it; on the View side so do RecyclerView and ListView.

findNodes(testTag: "feed")             → { "nodes": [{ "id": 12, "actions": ["ScrollBy", "ScrollToIndex", …] }] }
performNodeAction(nodeId: 12, action: "ScrollToIndex", index: 240)
findNodes(text: "Item 240")            → now composed, and BringIntoView can finish the job if needed

Both scrolls are applied by the container on its next frame, so a capture taken in the same breath still shows the old bounds — capture again after.

com.kitakkun.jetwhale.semantics.getViewAttributes ​

Reads one Android View node's platform attributes, addressed by rootId and nodeId. Each entry carries its id (what setViewAttribute names), label, group, type, value as a string, the options of an enum, and editable: false when it cannot be written. A node that has no attributes — a Compose node — comes back as a message rather than an error. See Editing View attributes.

type is one of bool, int, float, text, color, dimension, enum and layoutSize, and it is the whole of what the attribute takes: a dimension is a pixel figure and nothing else, an enum is one of its options and nothing else. A layoutSize — layout.width, layout.height — is the one that takes either, and it lists its constants whichever it currently reads as, so both possibilities are visible from a single read:

{ "id": "layout.width", "type": "layoutSize", "value": "WRAP_CONTENT",
  "constants": ["MATCH_PARENT", "WRAP_CONTENT"] }
{ "id": "layout.width", "type": "layoutSize", "value": "500.0", "dp": 250.0,
  "constants": ["MATCH_PARENT", "WRAP_CONTENT"] }

com.kitakkun.jetwhale.semantics.setViewAttribute ​

Changes one attribute: rootId, nodeId, attributeId, and value as a string, read according to the attribute's own type — "GONE", "true", "0.5", "#80FF0000", "24" — so there is no sealed JSON to construct. A layoutSize takes one of its constants, case-insensitively, or a pixel figure: "wrap_content", "match_parent", "500". The answer carries the attribute as it reads back afterwards, which is not always what was asked for: an app may clamp a value or ignore it. The edit is temporary, and only View nodes have attributes at all.

findNodes(resourceId: "status")   → { "nodes": [{ "rootId": "android-window-1f2e", "id": -4, "kind": "View" }] }
getViewAttributes(rootId: "android-window-1f2e", nodeId: -4)
setViewAttribute(rootId: "android-window-1f2e", nodeId: -4, attributeId: "textColor", value: "#FF0000FF")

Why not the CLI? ​

android layout and adb shell uiautomator dump both go out to the accessibility framework across a process boundary and write a file on the device before anything can read it. This plugin reads the semantics tree inside the app, on its main thread, and sends it back over the JetWhale connection that is already open — so a capture costs about as much as a frame.

Measured on one machine, on the same screen, back to back — a Pixel-class emulator (1080×2400, density 2.625) showing the demo app's Compose nodes screen, 38–39 elements:

median
com.kitakkun.jetwhale.semantics.getNodeTree (host → app → host)14 ms (min 12, p90 15)
⤷ of which reading the tree on the device1 ms (max 6)
com.kitakkun.jetwhale.semantics.findNodes11 ms
adb shell uiautomator dump + adb pull1,960 ms
android layout (Google's Android CLI)2,703 ms

That is ~190× faster than android layout on this setup — fast enough that an agent can capture between every action instead of budgeting for the dump. Treat the ratio as indicative rather than a spec: an emulator is slower than a physical device, and a bigger screen means more nodes.

The host shows both numbers live — what the capture cost on the device, and the round trip from the host — so a slow capture says where the time went.

The tree is also richer than what the CLIs return: android layout gives a flat list with text, content-desc and bounds, but no testTag, no role and no per-node id, so an agent can only aim by label or by pixel. This plugin reports all three, which is what makes performNodeAction able to address a node directly.

Are the coordinates right? ​

bounds and tap are screen coordinates in the node's unit (pixels on Android and desktop, points on iOS, which is what idb ui tap takes), so they have to survive whatever the device does to the window. They were checked two ways at once — cross-checked against android layout's reading of the same screen, and proved by tapping the reported point and watching the intended node react — across the conditions that move a window around:

Conditionnodes cross-checkedworst disagreementtap reached the node
gesture nav, portrait, density 420191 px✅
3-button navigation bar181 px✅
landscape101 px✅
density 320291 px✅
800×1280 @ density 320121 px✅

The residual 1 px is rounding: this plugin rounds, android layout truncates.

Each condition was also re-run with a dialog open, which is the case that actually exercises the arithmetic — a dialog is its own window and does not start at the screen origin, so a window-relative coordinate reported as if it were absolute would be off by the whole offset. In landscape that offset reaches (717, 298); the dialog's nodes still agreed to within 1 px, and tapping the reported point closed the dialog. Every root reports its own windowOffset, so a suspicious coordinate can be traced back to the window it came from.

Platform support ​

The capture and action layer is written against SemanticsOwner, which lives in Compose's common source set — so it is the same code on every target. What differs is only how a probe finds an owner, and that is where platform support begins and ends.

TargetProbeWhat you write
Android✅installJetWhaleSemanticsProbe(application), or JetWhaleSemanticsProbe() in a composition — captures the window, Android Views included
Desktop (JVM)✅JetWhaleSemanticsProbe() inside your Window { }, under @OptIn(ExperimentalComposeUiApi::class)
iOS✅installJetWhaleSemanticsProbe() at startup — reads the accessibility tree, so UIKit and SwiftUI come along; see iOS support
JS, Wasm—the standard entry points expose no owner — see Web

Those are the targets the agent artifact ships for — Compose Multiplatform's own set. Linux, mingw and macOS are absent: the first two have no androidx.compose.ui at all, and macOS needs the whole build to opt into Compose's experimental support for it.

Android finds roots process-wide through the callback Compose fires when it creates the view backing a composition. Desktop has no such callback, so its probe is scoped to a window and reads ComposeWindow.semanticsOwners — which is snapshot-backed, so a dialog or popup rendered inside that window appears and disappears on its own. A second Window { } is a second composition and needs its own call. ComposeWindow.semanticsOwners is @ExperimentalComposeUiApi, so the desktop probe carries that marker too — opt in at the call site (@OptIn(ExperimentalComposeUiApi::class)), or the module will not compile. The Android probe has no such requirement.

iOS is the odd one out: it has a probe, but not one that reads a SemanticsOwner. Compose Multiplatform hands none out for the scene an app shows — ComposeUIViewController builds its ComposeScene internally, and PlatformContext.SemanticsOwnerListener, the seam that would deliver the owners, is only consulted for a scene the caller constructs. What Compose does publish on iOS is its accessibility tree, and that is what the iOS probe reads — see iOS support for what that changes.

Web ​

There is no probe on JS or Wasm: ComposeViewport builds its scene internally, like iOS, and the browser has no accessibility tree the agent could read instead. Desktop was in the same position until Compose Multiplatform 1.10 exposed ComposeWindow.semanticsOwners; web has no equivalent yet. The plugin's capture and action layer is already common code, so a probe is a small addition once an owner can be reached.

Until then, registerSemanticsOwner is the seam to use for any host you build yourself on top of ComposeScene — including ImageComposeScene, whose semanticsOwners is available on these targets:

kotlin
import com.kitakkun.jetwhale.plugins.semantics.agent.ComposeNodeSourceRegistry
import com.kitakkun.jetwhale.plugins.semantics.agent.registerSemanticsOwner

ComposeNodeSourceRegistry.registerSemanticsOwner(
    owner = mySemanticsOwner,
    sourceId = "main-window",
    label = "Main window",
    density = 2f,
)

Which thread reads the tree

Semantics may only be read on the thread that owns the composition, and the probes get there without adding a dependency to your app — a Handler on Android, EventQueue on desktop. A target without a probe falls back to Dispatchers.Main, which on desktop would need kotlinx-coroutines-swing on your classpath; pass your own ComposeUiThread to registerSemanticsOwner if that does not suit.

Troubleshooting ​

"No root is registered." No probe is installed. Add installJetWhaleSemanticsProbe(application) on Android, installJetWhaleSemanticsProbe() on iOS, or JetWhaleSemanticsProbe() inside your composition. On JS and Wasm this is expected — see Web.

A dialog's contents are missing. On Android a dialog is a separate window. The Application-level probe finds it; an in-composition probe only registers the window it was called in. A dialog built from Android views with no Compose in it is not captured at all — see Android View support. On iOS a dialog stays inside the window that showed it and is part of that root.

An action comes back performed: false. The message says why — the node does not expose that action, it is disabled, or its handler declined. Capture the tree again and check the node's actions list; only what is listed there can be invoked.

Node ids changed between calls. A node's id is stable while it stays composed — a View node's, while the view is alive — and ids are only unique within their root. After anything that recomposes the screen, capture again rather than reusing ids.

Released under the Apache License 2.0.