genroc

docs / Guides / Process definition

Error handling

Learn how to deal with failures

You might be wondering what happens when a fetch action fails. Let’s look at that.

  - id: load_user
    action:
      type: fetch
      method: get
      url: "https://api.example.com/users/${input.user_id}"
      responses:
        200:
          type: object
          properties:
            email: { type: string }
            activated: { type: boolean }
          required: [email, activated]
        409:
          type: object
          properties:
            user_is_protected: { type: boolean }
      timeout: 10s   # default is 30s
    on_error:
      - code: ["http.409"]
        case: "error.data.user_is_protected == true"
        # a known error from the server
        goto: $user-protected
      - code: ["pre.%", "http.timeout", "http.disconnected", "http.503"]
        # server is down, or the connection was aborted
        retry: 3
    switch: next

Your main tool is the on_error field. It is evaluated from top to bottom, similarly to switch. You can match one or multiple errors with the code field and you can also use % wildcard to match multiple errors at once.

The error object

The error object has non-nullable code, message and task (containing the task name).

You can type your error responses the same way as your success responses. When you do that, you can match the typed error response code and then read the error response body through error.data.

If you match multiple error codes in code, error.data widens to cover all of them, so it is better to always match errors that carry the same data type.

Responses field

There are multiple ways to define the responses object.

responses:
  "200,201":          # comma separated list
    type: object
    properties:
      success: ...
  4xx:                # x matches any number
    type: object
    properties:
      error: ...

Case field

If you want to perform an additional check on the error response body, you can use the case field, which works the same way as on switch, so you can write a custom expression returning a boolean.

You can also use the case field alone (without matching the code), but for most cases the code field is more convenient.

Both selectors have to hold. A rule whose code matches but whose case is false does not catch the error - it falls through to the next rule, so you can write several rules for one code and let them narrow each other.

Custom timeout

The default timeout is 30s, but you can specify your own. It is also a slot (static values can be human-readable, dynamic ones milliseconds only). If the fetch doesn’t respond in time you will receive http.timeout, or pre.timeout if the request had not been sent yet.

Reacting to the error

To react to an error you can retry or use goto. You can also combine them: the task is retried first, and when the retries are depleted it follows the goto.

Retry

This parks the process for some time and then tries the fetch again. There are advanced settings for retry as well:

retry: 3
# or
retry:
  retries: 3
  delay: "1s"
  factor: 2
  max_delay: "5m"
  • retries - the extra fetches beyond the first one. So 3 retries can give you 4 failed fetches.
  • delay - how long to wait before the next retry (1s is default)*
  • factor - a multiplier of the delay after each retry (1s, 2s, 4s in this case, 2 is default)
  • max_delay - a cap on the multiplication (max(5m, delay) is default)

All the fields are slots whose context contains error, so you can compute the delay from the response body. The delay and max_delay slots accept only a number of milliseconds when dynamic.

* the actual wait is randomised between half the computed delay and the full one, so a fleet doesn’t hammer a recovering endpoint in lockstep; jitter only shortens, never exceeds max_delay

Goto

If you want the process to continue, you can use the goto field to route to another task (or end the process successfully with end). The targeted task will have an extra last_error field available, with the error data. The field is available only to the task directly after the error; later tasks have no access to it.

  - id: load_user
    action:
      type: fetch
      ...
    on_error:
      - code: ["http.409"]
        goto: $user-protected
    ...
  - id: user-protected
    # you can access the `last_error` variable
    output:
      is_protected: "$: last_error.data.user_is_protected"
    ...

Panic

An error no rule handles - or one that runs out of retries - fails the process on its own. The state switches to failed, execution stops, and error_code is the error’s own code (http.500). Failed is a retryable state, so the process can be retried with the genctl retry command.

panic is that same ending with a code you choose instead. Since error_code is what you filter and alert on, a panic is how a failure gets a name from your domain rather than from the transport. You can call it from on_error and switch blocks.

  on_error:
    - code: ["http.401"]
      panic:
        code: "not_authenticated"
        message: "Not authenticated to access the API"
        data: <optional data>

Raise

Raise is similar to panic, but it sends the process into the non-retryable raised state. A process which raised an error cannot be retried.

  on_error:
    - code: ["http.404"]
      raise:
        code: "item_not_found"
        message: "The ${input.item_id} item was not found"
        data: <optional data>

The raised state will make more sense once we get to child processes later.

Only once

In certain cases it is important that an operation happens exactly once (e.g. money transfers).

Genroc has a primitive for this, called only_once.

  - id: transfer_money
    action:
      type: fetch
      method: post
      url: "https://api.example.com/money/transfer"
      body:
        from: "$: input.from"
        to: "$: input.to"
        amount: "$: input.amount"
    only_once: true
    on_error:
      - code: ["pre.%"]
        retry: 3
      - code: ["http.503"]
        retry: 3
        not_reached: true
    switch: end

In this example we are using only_once, so Genroc is careful to hit the endpoint only once. You can safely retry on pre.% errors, because those errors happen before the request leaves (DNS lookup, TCP handshake).

You can also retry on other errors, but you have to explicitly add the not_reached field. That field is there to make sure you don’t retry by accident: Genroc requires it because you have to be certain the error really means the operation was not performed. You cannot use wildcards on codes when not_reached is used.

If a task marked only_once fails, genctl retry will refuse. If you really want to retry, use genctl retry --force.

Ambiguous cases

There are cases where we genuinely cannot know whether the operation was performed (read about the Two Generals’ Problem). In these cases you need to check through a different channel whether the money was transferred.

CodeReported byWhat happened
only_once.interruptedany only_once taska previous attempt was interrupted - route it, never retry
http.timeoutfetchconnected, but no response arrived in time
http.disconnectedfetchthe request went out, the connection broke before a response
external.timeoutexternalthe wait deadline elapsed
external.lostexternala worker’s claim expired without an answer

not_reached buys nothing here: it asserts what an error means, and that is a claim you can only make about an error that came back. Naming one of these alongside retry is refused at registration.

What to do with ambiguous cases

By default the process will fail, and you can manually check whether the money went out.

Or you can have an idempotent task that checks whether the transaction happened, and retry based on that.

  - id: transfer_money
    action:
      type: fetch
      method: post
      url: "https://api.example.com/money/transfer"
      body:
        from: "$: input.from"
        to: "$: input.to"
        amount: "$: input.amount"
        transaction_id: "$: input.transaction_id"
    only_once: true
    on_error:
      - code: ["only_once.interrupted", "http.timeout", "http.disconnected"]
        goto: $check_transfer
    switch: end
  - id: check_transfer
    action:
      type: fetch
      method: get
      url: "https://api.example.com/money/check_transaction/${input.transaction_id}"
      accepted_status: [2xx, "404"]   # which statuses are not considered an error
    output:
      retries: "$: (self.previous.retries ?? 0) + 1"
    switch:
      - case: "self.status == 404 && self.output.retries < 3"
        goto: $transfer_money
      - goto: end

This is a slightly more complicated example, but what we do is check in the check_transfer task whether the transaction exists. We limit the retries to 3 through the output mapping (we read the previous output.retries and add one).

Accepted statuses

In the example above, we’ve introduced a new field accepted_status. This field defines which HTTP statuses are considered a success.

The default value is 2xx, but there is a caveat when the responses field defines an output.

Responses field influence

If the responses field contains a 2xx key, the accepted statuses are narrowed to only those.

This can feel a bit strange, because if you define a response type only for 200, you can receive 202 as an error. The logic is that once you have declared a response type, you usually need the data from the body.

If you need to always accept all 2xx statuses regardless of the body, you can override accepted_status manually. But you then have to account for the case where the body won’t be there.

Fetch error codes

Errors are sorted into families by their prefix:

CodeMeaning
http.<status>a status accepted_status did not admit, e.g. http.404
http.timeoutconnected, but no response arrived in time
http.disconnectedthe request went out, the connection broke before a response
pre.timeouttimed out before the request was written - it never left
pre.errorfailed before the request was written - it never left
result.parsethe response body was not valid JSON
result.too_largethe response body exceeded the size a fetch will read
result.invalidthe result did not satisfy the schema declared for it
only_once.interrupteda worker didn’t finish an only_once task; a regular task would simply start over, but this one cannot