Lectures

Senior

Coroutines Deep Dive

Job lifecycle, failure propagation, supervision, cooperative cancellation, dispatchers, and shared-state reasoning.

coroutinesjobssupervisiondispatchersconcurrency
On this page

Predict coroutine lifecycle and concurrency behavior from ownership and context. All supplied examples are preserved as read-only fragments; none was compiled or executed during preparation.

Learning objectives:

  • distinguish body completion, Job completion, cancellation, and failure;
  • place supervision at the right sibling boundary;
  • reason about dispatcher views, thread-local propagation, and shared state;
  • explain interview cases from their Job trees.

Theory

Job lifecycle

Job is a coroutine lifecycle, not merely a cancel handle. Useful states include New, Active, Completing, Cancelling, and Completed. isActive, isCompleted, and isCancelled describe different lifecycle properties; completion of the body and completion of the Job are not necessarily simultaneous.

           start
NEW ───────────────► ACTIVE
                       │
             ┌─────────┴─────────┐
             │                   │
          success             failure /
             │                cancellation
             ▼                   │
        COMPLETING          CANCELLING
             │                   │
             └─────────┬─────────┘
                       ▼
                   COMPLETED
job.isActive
job.isCompleted
job.isCancelled

Official reference: Job API.

Why Completing exists

The parent body can return while its Job still waits for children. The outer Job completes only after its nested child settles, so structured lifetime extends beyond the last statement in the parent body.

val job = launch {
    launch {
        delay(5000)
    }

    println("parent body finished")
}
Parent body:
█████

Child:
██████████████████████████

Parent Job:
██████████████████████████

Cancellation travels down

Canceling the parent recursively cancels its children. Trace cancellation along the actual Job hierarchy rather than treating launches as unrelated background tasks.

val parent = scope.launch {
    launch { delay(10_000) }
    launch { delay(10_000) }
}

parent.cancel()
parent.cancel()
      │
      ▼
 Parent cancelled
      │
 ┌────┴────┐
 ▼         ▼
Child A   Child B
cancel    cancel

Official reference: Cancellation guide.

Child failure travels up

An uncaught non-cancellation failure in a normal child fails its parent; the failed parent cancels its other children. Failure travels upward and the resulting cancellation travels downward.

scope.launch {
    launch {
        delay(100)
        error("BOOM")
    }

    launch {
        delay(10_000)
        println("finished")
    }
}
Child A failure
      ↓
    Parent
      ↓
 Child B cancellation

Official reference: Exception propagation.

CancellationException is not ordinary failure

RuntimeException represents failure; CancellationException represents cancellation. Do not convert normal cancellation into a business error or silently swallow it.

RuntimeException
       ↓
     failure
       ↓
cancel parent

CancellationException
       ↓
 cancellation

Cooperative cancellation

A compute loop without cancellable suspension points or explicit checks can continue after cancel. Use isActive for a normal loop exit or ensureActive to throw on cancellation.

val job = launch(Dispatchers.Default) {
    while (true) {
        calculateSomething()
    }
}

job.cancel()
while (isActive) {
    calculate()
}
while (true) {
    ensureActive()
    calculate()
}

isActive and ensureActive

isActive lets code choose how to exit. ensureActive throws CancellationException when canceled, preserving standard cancellation semantics.

Official reference: ensureActive API.

yield

yield checks cancellation and gives the scheduler an opportunity to execute other coroutines. It should not conceal an unbounded or poorly designed CPU task.

while (true) {
    processChunk()
    yield()
}

Official reference: yield API.

Cleanup in finally

finally runs during cancellation. A cancellable suspending cleanup inside it may immediately observe the canceled context again; design cleanup according to that requirement.

try {
    doWork()
} finally {
    cleanup()
}

NonCancellable

A short suspending cleanup that must finish can use withContext(NonCancellable). Do not turn this into a large network synchronization or long business operation that ignores cancellation.

finally {
    withContext(NonCancellable) {
        saveState()
    }
}

Official reference: NonCancellable API.

Why an outer catch misses child failure

A try/catch around launch surrounds coroutine creation; the child exception occurs later during its execution. Catch inside the coroutine or at an appropriate structured boundary.

try {
    scope.launch {
        throw RuntimeException("Boom")
    }
} catch (e: Exception) {
    println("Caught")
}

async failure

await throws the failure retained by Deferred, but failure can already affect its parent hierarchy. A failing child async can cancel a normal coroutineScope before the caller reaches await.

val deferred = scope.async {
    error("Boom")
}
coroutineScope {
    val deferred = async {
        error("Boom")
    }

    delay(10_000)
    deferred.await()
}

Official reference: Deferred API.

A fail-fast result boundary

When every component is required for one Page, coroutineScope with independent async children expresses a single fail-fast result. A child failure cancels siblings and is propagated to the caller after child cleanup.

suspend fun loadPage(): Page =
    coroutineScope {
        val user = async { loadUser() }
        val posts = async { loadPosts() }

        Page(
            user.await(),
            posts.await(),
        )
    }

Official reference: coroutineScope API.

Independent sibling work

supervisorScope is suitable when each sibling result is useful independently. A failed advertising request need not cancel a profile request, but the failure still needs a handling or reporting policy.

supervisorScope {
    launch { loadProfile() }
    launch { loadAds() }
}

Official reference: supervisorScope API.

The SupervisorJob trap

A SupervisorJob above one ordinary outer launch does not supervise that launch’s grandchildren. A and B still share the ordinary outer parent: A can fail it and cancel B. Place supervisorScope at the sibling boundary that needs isolation.

val scope = CoroutineScope(
    SupervisorJob() + Dispatchers.Default
)

scope.launch {
    launch { taskA() }
    launch { taskB() }
}
SupervisorJob
      │
      ▼
 outer launch
      │
   ┌──┴──┐
   ▼     ▼
   A     B
A fails
  ↓
outer launch fails
  ↓
B cancelled

Business errors and exception handlers

CoroutineExceptionHandler handles uncaught exceptions at a suitable root boundary. Map expected IOException or domain errors locally; a handler is not a substitute for business-error control flow.

try {
    repository.load()
} catch (e: IOException) {
    // map to UI/domain error
}

Official reference: CoroutineExceptionHandler API.

Read-modify-write races

counter++ reads, increments, and writes. Two concurrent coroutines can read 10 and both write 11, losing one increment. Coroutines themselves do not make shared mutation atomic.

var counter = 0

coroutineScope {
    repeat(1000) {
        launch(Dispatchers.Default) {
            counter++
        }
    }
}
Thread A          Thread B

read 10           read 10
+1                +1
write 11          write 11

Mutex and alternatives

Mutex can protect a critical invariant, but it is not the automatic answer. Alternatives include immutable state, atomics, MutableStateFlow.update, Channel or actor-like serialization, and thread confinement.

val mutex = Mutex()

mutex.withLock {
    counter++
}

Official reference: Shared mutable state.

Thread confinement

One state owner can serialize mutation instead of scattering locks. Events reach the owner, which updates state. This removes many races through ownership and sequencing.

Event A ──┐
Event B ──┼──► State owner ──► State
Event C ──┘

Default dispatcher

Use Default for CPU work such as move generation, game-tree search, evaluation, sorting, compression, parsing, and cryptography.

withContext(Dispatchers.Default) {
    calculateMoves()
}

Official reference: Dispatchers API.

IO dispatcher

Use IO for genuinely blocking I/O such as legacy file, socket, or JDBC APIs. The fact that an operation involves a network does not alone mean it blocks.

withContext(Dispatchers.IO) {
    legacyBlockingApi()
}

limitedParallelism

A dispatcher view can limit parallel execution for a subsystem, such as an IO view with four workers or a CPU view with two. This limits simultaneously executing tasks, not necessarily the number of suspended operations or in-flight requests; use a separate permit mechanism for that invariant.

val dispatcher =
    Dispatchers.IO.limitedParallelism(4)
val botDispatcher =
    Dispatchers.Default.limitedParallelism(2)

Official reference: limitedParallelism API.

Unconfined dispatcher

Unconfined begins in the current call stack and thread. After suspension, it can resume where the suspending operation resumes it. Ordinary Android application code rarely needs this behavior.

Official reference: Context and dispatchers.

A dispatcher is not a new thread

withContext(IO) does not mean new Thread(). A dispatcher expresses scheduling policy, and its implementation may reuse or share threads.

withContext(Dispatchers.IO) { ... }

Thread identity can change

A coroutine can execute on different physical threads before and after suspension while preserving its context contract. Do not base correctness on Thread.currentThread identity.

ThreadLocal and context propagation

ThreadLocal belongs to a physical thread, while coroutines can migrate. For explicit thread-local propagation, asContextElement supplies a coroutine context element; it does not make arbitrary thread-local writes automatically trackable.

Official reference: asContextElement API.

Practice

Trace the supplied cases, identify the owner and failure boundary, and explain the observations before running an experiment.

Interview: an ordinary scope

Purpose: trace failure and cleanup. A throws, the ordinary parent scope fails, B is canceled, B’s finally runs, and the failure propagates to the caller. Explain the causal chain rather than promising arbitrary concurrent print ordering.

coroutineScope {
    launch {
        delay(100)
        throw RuntimeException("A")
    }

    launch {
        try {
            delay(10_000)
            println("B completed")
        } finally {
            println("B finally")
        }
    }
}
A throws
   ↓
parent scope fails
   ↓
B cancelled
   ↓
B finally executes
   ↓
RuntimeException propagates to caller

Interview: a supervisor scope

Purpose: trace independent failure. A’s failure does not cancel B, so B can complete. Explain where A’s uncaught exception is reported or handled; supervision alone does not swallow it.

supervisorScope {
    launch {
        delay(100)
        throw RuntimeException("A")
    }

    launch {
        delay(500)
        println("B completed")
    }
}

Interview: an unnecessary nested launch

Purpose: find a mistaken catch boundary. The outer try does not catch a later asynchronous child failure as expected. Often remove the nested launch and call the suspending repository function directly, explicitly rethrowing CancellationException before mapping other errors.

viewModelScope.launch {
    try {
        launch {
            repository.load()
        }
    } catch (e: Exception) {
        showError()
    }
}
viewModelScope.launch {
    try {
        repository.load()
    } catch (e: CancellationException) {
        throw e
    } catch (e: Exception) {
        showError()
    }
}

A coroutine analysis algorithm

Draw the Job hierarchy; identify every child’s parent; find supervision boundaries; locate failure or cancellation; trace failure upward and cancellation downward; only then analyze try/catch, await, and the exception handler.

Further reading