Senior
Coroutines Deep Dive
Job lifecycle, failure propagation, supervision, cooperative cancellation, dispatchers, and shared-state reasoning.
On this page
Predict coroutine lifecycle and concurrency behavior from ownership and context. All supplied examples are preserved as read-only fragments; none was compiled or executed during preparation.
Learning objectives:
- distinguish body completion, Job completion, cancellation, and failure;
- place supervision at the right sibling boundary;
- reason about dispatcher views, thread-local propagation, and shared state;
- explain interview cases from their Job trees.
Theory
Job lifecycle
Job is a coroutine lifecycle, not merely a cancel handle. Useful states include New, Active, Completing, Cancelling, and Completed. isActive, isCompleted, and isCancelled describe different lifecycle properties; completion of the body and completion of the Job are not necessarily simultaneous.
start
NEW ───────────────► ACTIVE
│
┌─────────┴─────────┐
│ │
success failure /
│ cancellation
▼ │
COMPLETING CANCELLING
│ │
└─────────┬─────────┘
▼
COMPLETED
job.isActive
job.isCompleted
job.isCancelled
Official reference: Job API.
Why Completing exists
The parent body can return while its Job still waits for children. The outer Job completes only after its nested child settles, so structured lifetime extends beyond the last statement in the parent body.
val job = launch {
launch {
delay(5000)
}
println("parent body finished")
}
Parent body:
█████
Child:
██████████████████████████
Parent Job:
██████████████████████████
Cancellation travels down
Canceling the parent recursively cancels its children. Trace cancellation along the actual Job hierarchy rather than treating launches as unrelated background tasks.
val parent = scope.launch {
launch { delay(10_000) }
launch { delay(10_000) }
}
parent.cancel()
parent.cancel()
│
▼
Parent cancelled
│
┌────┴────┐
▼ ▼
Child A Child B
cancel cancel
Official reference: Cancellation guide.
Child failure travels up
An uncaught non-cancellation failure in a normal child fails its parent; the failed parent cancels its other children. Failure travels upward and the resulting cancellation travels downward.
scope.launch {
launch {
delay(100)
error("BOOM")
}
launch {
delay(10_000)
println("finished")
}
}
Child A failure
↓
Parent
↓
Child B cancellation
Official reference: Exception propagation.
CancellationException is not ordinary failure
RuntimeException represents failure; CancellationException represents cancellation. Do not convert normal cancellation into a business error or silently swallow it.
RuntimeException
↓
failure
↓
cancel parent
CancellationException
↓
cancellation
Cooperative cancellation
A compute loop without cancellable suspension points or explicit checks can continue after cancel. Use isActive for a normal loop exit or ensureActive to throw on cancellation.
val job = launch(Dispatchers.Default) {
while (true) {
calculateSomething()
}
}
job.cancel()
while (isActive) {
calculate()
}
while (true) {
ensureActive()
calculate()
}
isActive and ensureActive
isActive lets code choose how to exit. ensureActive throws CancellationException when canceled, preserving standard cancellation semantics.
Official reference: ensureActive API.
yield
yield checks cancellation and gives the scheduler an opportunity to execute other coroutines. It should not conceal an unbounded or poorly designed CPU task.
while (true) {
processChunk()
yield()
}
Official reference: yield API.
Cleanup in finally
finally runs during cancellation. A cancellable suspending cleanup inside it may immediately observe the canceled context again; design cleanup according to that requirement.
try {
doWork()
} finally {
cleanup()
}
NonCancellable
A short suspending cleanup that must finish can use withContext(NonCancellable). Do not turn this into a large network synchronization or long business operation that ignores cancellation.
finally {
withContext(NonCancellable) {
saveState()
}
}
Official reference: NonCancellable API.
Why an outer catch misses child failure
A try/catch around launch surrounds coroutine creation; the child exception occurs later during its execution. Catch inside the coroutine or at an appropriate structured boundary.
try {
scope.launch {
throw RuntimeException("Boom")
}
} catch (e: Exception) {
println("Caught")
}
async failure
await throws the failure retained by Deferred, but failure can already affect its parent hierarchy. A failing child async can cancel a normal coroutineScope before the caller reaches await.
val deferred = scope.async {
error("Boom")
}
coroutineScope {
val deferred = async {
error("Boom")
}
delay(10_000)
deferred.await()
}
Official reference: Deferred API.
A fail-fast result boundary
When every component is required for one Page, coroutineScope with independent async children expresses a single fail-fast result. A child failure cancels siblings and is propagated to the caller after child cleanup.
suspend fun loadPage(): Page =
coroutineScope {
val user = async { loadUser() }
val posts = async { loadPosts() }
Page(
user.await(),
posts.await(),
)
}
Official reference: coroutineScope API.
Independent sibling work
supervisorScope is suitable when each sibling result is useful independently. A failed advertising request need not cancel a profile request, but the failure still needs a handling or reporting policy.
supervisorScope {
launch { loadProfile() }
launch { loadAds() }
}
Official reference: supervisorScope API.
The SupervisorJob trap
A SupervisorJob above one ordinary outer launch does not supervise that launch’s grandchildren. A and B still share the ordinary outer parent: A can fail it and cancel B. Place supervisorScope at the sibling boundary that needs isolation.
val scope = CoroutineScope(
SupervisorJob() + Dispatchers.Default
)
scope.launch {
launch { taskA() }
launch { taskB() }
}
SupervisorJob
│
▼
outer launch
│
┌──┴──┐
▼ ▼
A B
A fails
↓
outer launch fails
↓
B cancelled
Business errors and exception handlers
CoroutineExceptionHandler handles uncaught exceptions at a suitable root boundary. Map expected IOException or domain errors locally; a handler is not a substitute for business-error control flow.
try {
repository.load()
} catch (e: IOException) {
// map to UI/domain error
}
Official reference: CoroutineExceptionHandler API.
Read-modify-write races
counter++ reads, increments, and writes. Two concurrent coroutines can read 10 and both write 11, losing one increment. Coroutines themselves do not make shared mutation atomic.
var counter = 0
coroutineScope {
repeat(1000) {
launch(Dispatchers.Default) {
counter++
}
}
}
Thread A Thread B
read 10 read 10
+1 +1
write 11 write 11
Mutex and alternatives
Mutex can protect a critical invariant, but it is not the automatic answer. Alternatives include immutable state, atomics, MutableStateFlow.update, Channel or actor-like serialization, and thread confinement.
val mutex = Mutex()
mutex.withLock {
counter++
}
Official reference: Shared mutable state.
Thread confinement
One state owner can serialize mutation instead of scattering locks. Events reach the owner, which updates state. This removes many races through ownership and sequencing.
Event A ──┐
Event B ──┼──► State owner ──► State
Event C ──┘
Default dispatcher
Use Default for CPU work such as move generation, game-tree search, evaluation, sorting, compression, parsing, and cryptography.
withContext(Dispatchers.Default) {
calculateMoves()
}
Official reference: Dispatchers API.
IO dispatcher
Use IO for genuinely blocking I/O such as legacy file, socket, or JDBC APIs. The fact that an operation involves a network does not alone mean it blocks.
withContext(Dispatchers.IO) {
legacyBlockingApi()
}
limitedParallelism
A dispatcher view can limit parallel execution for a subsystem, such as an IO view with four workers or a CPU view with two. This limits simultaneously executing tasks, not necessarily the number of suspended operations or in-flight requests; use a separate permit mechanism for that invariant.
val dispatcher =
Dispatchers.IO.limitedParallelism(4)
val botDispatcher =
Dispatchers.Default.limitedParallelism(2)
Official reference: limitedParallelism API.
Unconfined dispatcher
Unconfined begins in the current call stack and thread. After suspension, it can resume where the suspending operation resumes it. Ordinary Android application code rarely needs this behavior.
Official reference: Context and dispatchers.
A dispatcher is not a new thread
withContext(IO) does not mean new Thread(). A dispatcher expresses scheduling policy, and its implementation may reuse or share threads.
withContext(Dispatchers.IO) { ... }
Thread identity can change
A coroutine can execute on different physical threads before and after suspension while preserving its context contract. Do not base correctness on Thread.currentThread identity.
ThreadLocal and context propagation
ThreadLocal belongs to a physical thread, while coroutines can migrate. For explicit thread-local propagation, asContextElement supplies a coroutine context element; it does not make arbitrary thread-local writes automatically trackable.
Official reference: asContextElement API.
Practice
Trace the supplied cases, identify the owner and failure boundary, and explain the observations before running an experiment.
Interview: an ordinary scope
Purpose: trace failure and cleanup. A throws, the ordinary parent scope fails, B is canceled, B’s finally runs, and the failure propagates to the caller. Explain the causal chain rather than promising arbitrary concurrent print ordering.
coroutineScope {
launch {
delay(100)
throw RuntimeException("A")
}
launch {
try {
delay(10_000)
println("B completed")
} finally {
println("B finally")
}
}
}
A throws
↓
parent scope fails
↓
B cancelled
↓
B finally executes
↓
RuntimeException propagates to caller
Interview: a supervisor scope
Purpose: trace independent failure. A’s failure does not cancel B, so B can complete. Explain where A’s uncaught exception is reported or handled; supervision alone does not swallow it.
supervisorScope {
launch {
delay(100)
throw RuntimeException("A")
}
launch {
delay(500)
println("B completed")
}
}
Interview: an unnecessary nested launch
Purpose: find a mistaken catch boundary. The outer try does not catch a later asynchronous child failure as expected. Often remove the nested launch and call the suspending repository function directly, explicitly rethrowing CancellationException before mapping other errors.
viewModelScope.launch {
try {
launch {
repository.load()
}
} catch (e: Exception) {
showError()
}
}
viewModelScope.launch {
try {
repository.load()
} catch (e: CancellationException) {
throw e
} catch (e: Exception) {
showError()
}
}
A coroutine analysis algorithm
Draw the Job hierarchy; identify every child’s parent; find supervision boundaries; locate failure or cancellation; trace failure upward and cancellation downward; only then analyze try/catch, await, and the exception handler.