You might be wondering what happens when a fetch action fails. Let’s look at that.
- id: load_user
action:
type: fetch
method: get
url: "https://api.example.com/users/${input.user_id}"
responses:
200:
type: object
properties:
email: { type: string }
activated: { type: boolean }
required: [email, activated]
409:
type: object
properties:
user_is_protected: { type: boolean }
timeout: 10s # default is 30s
on_error:
- code: ["http.409"]
case: "error.data.user_is_protected == true"
# a known error from the server
goto: $user-protected
- code: ["pre.%", "http.timeout", "http.disconnected", "http.503"]
# server is down, or the connection was aborted
retry: 3
switch: next
Your main tool is the on_error field. It is evaluated from top to bottom, similarly to switch.
You can match one or multiple errors with the code field and you can also use % wildcard to match multiple errors at once.
The error object
The error object has non-nullable code, message and task (containing the task name).
You can type your error responses the same way as your success responses.
When you do that, you can match the typed error response code and then read the error response body through error.data.
If you match multiple error codes in
code,error.datawidens to cover all of them, so it is better to always match errors that carry the same data type.
Responses field
There are multiple ways to define the responses object.
responses:
"200,201": # comma separated list
type: object
properties:
success: ...
4xx: # x matches any number
type: object
properties:
error: ...
Case field
If you want to perform an additional check on the error response body, you can use the case field, which works the same way
as on switch, so you can write a custom expression returning a boolean.
You can also use the case field alone (without matching the code), but for most cases the code field is more convenient.
Both selectors have to hold. A rule whose code matches but whose case is false does not catch the error -
it falls through to the next rule, so you can write several rules for one code and let them narrow each other.
Custom timeout
The default timeout is 30s, but you can specify your own. It is also a slot (static values can be human-readable, dynamic ones milliseconds only).
If the fetch doesn’t respond in time you will receive http.timeout, or pre.timeout if the request had not been sent yet.
Reacting to the error
To react to an error you can retry or use goto. You can also combine
them: the task is retried first, and when the retries are depleted it
follows the goto.
Retry
This parks the process for some time and then tries the fetch again. There are advanced settings for retry as well:
retry: 3
# or
retry:
retries: 3
delay: "1s"
factor: 2
max_delay: "5m"
retries- the extra fetches beyond the first one. So 3 retries can give you 4 failed fetches.delay- how long to wait before the next retry (1sis default)*factor- a multiplier of the delay after each retry (1s,2s,4sin this case, 2 is default)max_delay- a cap on the multiplication (max(5m, delay) is default)
All the fields are slots whose context contains error, so you can
compute the delay from the response body.
The delay and max_delay slots accept only a number of milliseconds when dynamic.
* the actual wait is randomised between half the computed delay and the full one, so a fleet doesn’t hammer a recovering endpoint in lockstep; jitter only shortens, never exceeds max_delay
Goto
If you want the process to continue, you can use the goto field to route to another task (or end the process successfully with end).
The targeted task will have an extra last_error field available, with the error data. The field is available only to the
task directly after the error; later tasks have no access to it.
- id: load_user
action:
type: fetch
...
on_error:
- code: ["http.409"]
goto: $user-protected
...
- id: user-protected
# you can access the `last_error` variable
output:
is_protected: "$: last_error.data.user_is_protected"
...
Panic
An error no rule handles - or one that runs out of retries - fails the process on its own.
The state switches to failed, execution stops, and error_code is the error’s own code (http.500).
Failed is a retryable state, so the process can be retried with the genctl retry command.
panic is that same ending with a code you choose instead. Since error_code is what you filter
and alert on, a panic is how a failure gets a name from your domain rather than from the transport.
You can call it from on_error and switch blocks.
on_error:
- code: ["http.401"]
panic:
code: "not_authenticated"
message: "Not authenticated to access the API"
data: <optional data>
Raise
Raise is similar to panic, but it sends the process into the non-retryable raised state.
A process which raised an error cannot be retried.
on_error:
- code: ["http.404"]
raise:
code: "item_not_found"
message: "The ${input.item_id} item was not found"
data: <optional data>
The raised state will make more sense once we get to child processes later.
Only once
In certain cases it is important that an operation happens exactly once (e.g. money transfers).
Genroc has a primitive for this, called only_once.
- id: transfer_money
action:
type: fetch
method: post
url: "https://api.example.com/money/transfer"
body:
from: "$: input.from"
to: "$: input.to"
amount: "$: input.amount"
only_once: true
on_error:
- code: ["pre.%"]
retry: 3
- code: ["http.503"]
retry: 3
not_reached: true
switch: end
In this example we are using only_once, so Genroc is careful to hit the endpoint only once.
You can safely retry on pre.% errors, because those errors happen before the request leaves (DNS lookup, TCP handshake).
You can also retry on other errors, but you have to explicitly add the not_reached field.
That field is there to make sure you don’t retry by accident: Genroc requires it because you have to be
certain the error really means the operation was not performed. You cannot use wildcards on codes when not_reached is used.
If a task marked only_once fails, genctl retry will refuse. If you
really want to retry, use genctl retry --force.
Ambiguous cases
There are cases where we genuinely cannot know whether the operation was performed (read about the Two Generals’ Problem). In these cases you need to check through a different channel whether the money was transferred.
| Code | Reported by | What happened |
|---|---|---|
only_once.interrupted | any only_once task | a previous attempt was interrupted - route it, never retry |
http.timeout | fetch | connected, but no response arrived in time |
http.disconnected | fetch | the request went out, the connection broke before a response |
external.timeout | external | the wait deadline elapsed |
external.lost | external | a worker’s claim expired without an answer |
not_reached buys nothing here: it asserts what an error means, and that is a claim you can only
make about an error that came back. Naming one of these alongside retry is refused at registration.
What to do with ambiguous cases
By default the process will fail, and you can manually check whether the money went out.
Or you can have an idempotent task that checks whether the transaction happened, and retry based on that.
- id: transfer_money
action:
type: fetch
method: post
url: "https://api.example.com/money/transfer"
body:
from: "$: input.from"
to: "$: input.to"
amount: "$: input.amount"
transaction_id: "$: input.transaction_id"
only_once: true
on_error:
- code: ["only_once.interrupted", "http.timeout", "http.disconnected"]
goto: $check_transfer
switch: end
- id: check_transfer
action:
type: fetch
method: get
url: "https://api.example.com/money/check_transaction/${input.transaction_id}"
accepted_status: [2xx, "404"] # which statuses are not considered an error
output:
retries: "$: (self.previous.retries ?? 0) + 1"
switch:
- case: "self.status == 404 && self.output.retries < 3"
goto: $transfer_money
- goto: end
This is a slightly more complicated example, but what we do is check in the
check_transfer task whether the transaction exists.
We limit the retries to 3 through the output mapping (we read the previous output.retries and add one).
Accepted statuses
In the example above, we’ve introduced a new field accepted_status.
This field defines which HTTP statuses are considered a success.
The default value is 2xx, but there is a caveat when the responses field defines an output.
Responses field influence
If the responses field contains a 2xx key, the accepted statuses are narrowed to only those.
This can feel a bit strange, because if you define a response type only for 200, you can receive 202 as an error.
The logic is that once you have declared a response type, you usually need the data from the body.
If you need to always accept all 2xx statuses regardless of the body, you can override accepted_status manually.
But you then have to account for the case where the body won’t be there.
Fetch error codes
Errors are sorted into families by their prefix:
| Code | Meaning |
|---|---|
http.<status> | a status accepted_status did not admit, e.g. http.404 |
http.timeout | connected, but no response arrived in time |
http.disconnected | the request went out, the connection broke before a response |
pre.timeout | timed out before the request was written - it never left |
pre.error | failed before the request was written - it never left |
result.parse | the response body was not valid JSON |
result.too_large | the response body exceeded the size a fetch will read |
result.invalid | the result did not satisfy the schema declared for it |
only_once.interrupted | a worker didn’t finish an only_once task; a regular task would simply start over, but this one cannot |