Is appsheet being extremely slow today for everyone?

I will answer to the best of my knowledge:

  1. what is the actual issue? we know u said there is a latency observed during peak hours. pls give more details on this. is it the number of users for a single app (my app) that affects the app or the entire appsheet users thru out the world?

    The latency is a platform-wide infrastructure issue affecting AppSheet users globally, rather than something caused by your specific app or your specific user count. When global platform traffic surges during standard weekday business hours, this backend queue becomes backlogged, resulting in the extended write delays and sync timeouts you are experiencing.

  2. is the issue with the appsheet database? or the app?

    The issue is entirely on the AppSheet backend, It is not an issue with your specific app’s architecture, nor is it an issue with your underlying database/data source.

  3. what immediate steps can we (appsheet developers) take to improve the performance (I removed almost all the actions and multiple Virtuals columns and it still seems to show the red exclamation point during weekdays alone even when the edit is taking 5 secs in audit history)

    First, thank you for taking the initiative to optimize your application. While removing heavy Virtual Columns and complex actions is normally the best practice for improving standard app performance, it will not bypass the current issue. The root cause is on our end, and the fix must come from our engineering team.

  4. we understand you cannot promise when it would get resolved and that the team is also rolling out changes and then finding out it does not work. what exactly causes the delay to fix this problem.

    Fixing this has taken longer than expected because the bottleneck only exposes itself under massive, real-world global traffic. During off-peak hours and weekends, the system processes data normally and appears completely healthy.

    Because of this, engineering cannot easily simulate the failure in a test environment. They have had to write and deploy brand-new custom telemetry metrics directly into the production servers just to β€œsee” exactly which services are choking under pressure.

    The resolution requires an iterative cycle: the team deploys an optimization, waits for peak traffic to hit so the new metrics can record the results, and then uses that data to build the next iteration of the fix.

    We understand how frustrating this trial-and-error process is for your operations.

Thank you, Jose. I really appreciate your response!

Outstanding response, @Jose_Arteaga! Thank you!

it was worse in the morning.

just wanted to let you know so you can convey to the team.

hopefully gets better in the afternoon. thank you , Jose!

Following up on my post from Aug 20 with a week of fresh numbers, since @Jose_Arteaga said he’d pass our data to the bug. Short version: it has got materially worse this week, and our data independently confirms the weekend/weekday pattern Jose described in #79.

Same app as before β€” production field-services workforce management, AppSheet Database, us-east4, ~20 tables, 2,400 rows in the main table, 13–18 concurrent users on phones.


Failures per day

Sync / add / edit / delete operations terminating at or above 100 seconds, from the app’s own audit log. Bot rows and auth artifacts excluded.

Date Day Failures β‰₯100s Users hit
Aug 20 Thu 13 5 :yellow_circle: best weekday since the regression
Aug 21 Fri 39 9 :red_circle:
Aug 22 Sat 0 1 :green_circle: weekend
Aug 23 Sun 0 – :green_circle: weekend
Aug 24 Mon 84 9 :red_circle:
Aug 25 Tue 72 12 :red_circle:
Aug 26 Wed 104 9 :red_circle: still climbing at time of posting

Zero failures both weekend days. 84, 72 and 104 on the three working days either side of them.

@Rodrigo_Nuvens called this in #67 before it happened, and Jose confirmed the mechanism in #79. For anyone still trying to work out whether a given day’s improvement is a fix landing or just Saturday: in our data, weekends have been clean every single week since this started. Only weekday numbers are worth reading.


Latency β€” the part that doesn’t show up in failure counts

Today, Aug 26, 12:11–17:34 UTC, 14 active users:

Measure Aug 7 (a good day) Aug 26 (today)
Write operations 93 349
Median write 4.9 s 24.4 s
Max write 26.1 s 120.3 s
Median sync 4.4 s 15.2 s
Max sync 9.2 s 240.1 s
Ops over 100 s 0 119

The median is the number I’d draw attention to. A five-second save has become a twenty-four-second save. Most of that never appears as an error β€” it just makes the app unusable for people entering work on a phone in the field, and it doesn’t show up in any β€œfailure” metric.

240.1 s on a sync is the ~240 second sync ceiling. 120.3 s on a write is the row-operation ceiling. Both are being hit consistently again.


Error signatures, Aug 24–26 (333 failures)

Error Count
Data table '<X>' is not accessible… An error occurred while fetching rows from AppSheet database 152
N rows could not be updated in AppSheet database 37
A task was canceled. 33
Unable to find the specified row from AppSheet database 31
Timed out waiting for a resource 8
other / generic 72

That fifth one is the same string reported in #69 on Cloud SQL PostgreSQL. We are on AppSheet Database and getting it too. If the same β€œtimed out waiting for a resource” appears on Postgres and on ASDB, that points at something above the storage layer rather than at any one backend.


Which tables fail (Aug 24–26)

Table Times unreadable Rows in table
Subtasks 31 ~1,375
WorkOrderAttachments 30 –
Comments 21 2,047
Roles 19 7
SubtaskInventory 16 –
WorkOrderHistory 16 –
WorkOrderInventory 11 –
WorkOrders 8 2,410
SubtaskComments 6 –
Users 1 31

Repeating the point from my last post because it still holds: Roles has seven rows and two columns. It is a static lookup that never changes. The database failed to read it 19 times in three days. Ten distinct tables in one database, differing in size, schema and access pattern, all intermittently unreadable.


@Jose_Arteaga β€” thank you for #82. Having it stated plainly that this is β€œa platform-wide infrastructure issue affecting AppSheet users globally, rather than something caused by your specific app or your specific user count” is genuinely useful to those of us who build for clients. I have spent three weeks proving that to a customer from the outside; one sentence from you does it better.

If it helps engineering, happy to share full audit exports, request IDs or stack traces from any of the windows above.

Outstanding contribution, @Jonathan_Kitts!

Good news from our side for a change. Posting it because a clean-day datapoint is probably as useful to engineering as a bad one, and keeping the same format as my previous posts so it’s comparable.

Same app as before β€” production field-services workforce management, AppSheet Database, us-east4, ~20 tables, 13–18 concurrent users on phones.


This week, day by day

All five days sampled over the identical 13:00–17:00 UTC window (peak business hours for our crews). Same query, no row-cap truncation on any of them, so these are directly comparable.

Day Users Writes Median write Max write Syncs Median sync Max sync Ops >100s Failures
Mon Aug 24 9 132 32.5 s 153.3 s 92 15.2 s 157.3 s 30 19
Tue Aug 25 13 203 16.8 s 138.6 s 119 6.3 s 181.9 s 43 31
Wed Aug 26 13 309 38.5 s 120.3 s 77 30.1 s 175.6 s 114 107
Thu Aug 27 9 117 9.6 s 120.2 s 99 4.7 s 187.0 s 26 11
Fri Aug 28 10 99 8.5 s 32.7 s 135 4.0 s 7.6 s 0 0

Wednesday was the worst four hours we have recorded in this entire incident: a 38.5 second median write and a 30.1 second median sync, with 114 operations over 100 seconds.

Friday, same window, same app: zero failures, zero operations over 100 seconds, worst write 32.7 s and worst sync 7.6 s.

The maximums are the column I would draw attention to. Monday through Thursday, max sync sat between 157 and 187 seconds every single day. Friday it was 7.6 seconds. That is not a marginal improvement, the tail has simply gone.


Caveats, including against my own conclusion

  1. Friday had lower write volume β€” 99 writes against Wednesday’s 309. Load is the variable that exposes this (per @Jose_Arteaga in #79 and @Rodrigo_Nuvens in #67), so some of Friday’s improvement may be a lighter afternoon rather than a fix. Thursday is the more useful comparison: 117 writes and still a 9.6 s median with 26 slow operations, so Friday is better than Thursday on similar volume.

  2. We have been fooled twice already. Aug 6 looked like recovery and held four working days before collapsing to 103 failures. Aug 20 held exactly one day. Both times I thought it was over.

  3. Monday Aug 31 is the real test. Monday has been the worst day of every week since this started.

What feels different this time is that it is not a smaller number of timeouts, it is none, with nothing anywhere near either ceiling. A four-hour window where the slowest sync is under eight seconds is a different shape of result than a window with fewer 120-second failures.


@Jose_Arteaga β€” if something shipped Wednesday night or Thursday, our data is consistent with it working. Worth engineering knowing that one account went from 107 failures and a 38.5 s median write on Wednesday to zero failures and 8.5 s on Friday, in the same four-hour window.

Would others post whether Thursday and Friday improved for them?

Not calling it fixed. Just reporting a genuinely good day after three weeks of bad ones.

Yes , Thursday and Friday were much better! Really hoping it does not come back again on Monday!

Hi @Jose_Arteaga ,

We’re still experiencing the same slowness as before, at least on our end.

It doesn’t seem like anything has changed so far. Do you have any updates or new information regarding the issue?

Me too, syncing very slow!

seems like it means there was just no peak traffic on thursday and friday :confused:

Same here…

Monday follow-up to my Friday post. I said Monday would be the real test, so here it is: the improvement did not hold.

But the shape of today’s failure is different from last week’s, and I think that difference might actually be useful to engineering.


Friday vs Monday, identical window

Same app, same us-east4 database, same query, and I deliberately used the exact same 13:00–14:45 UTC window on both days so the volumes line up. Neither hit the row cap.

Measure Fri Aug 28 Mon Aug 31
Active users 9 10
Write operations 70 81
Median write 8.8 s 7.8 s
Max write 32.7 s 120.2 s
Sync operations 80 75
Median sync 4.0 s 4.5 s
Max sync 7.5 s 240.6 s
Operations over 100 s 0 28
Failures 0 22

Comparable users, comparable volume, ninety minutes each.


The part I think matters

The medians are unchanged. Monday’s median write is actually slightly faster than Friday’s β€” 7.8 s against 8.8 s. Median sync is 4.5 s against 4.0 s. By those numbers Monday looks like a healthy day.

Yet in the same window, 28 operations blew past 100 seconds and 22 failed outright, both ceilings hit cleanly (120.2 s on a write, 240.6 s on a sync).

This is not what last Wednesday looked like. On Aug 26 the median write was 38.5 s and the median sync 30.1 s β€” everything was slow, uniformly. Today, almost everything is fine and a subset of operations simply hangs until it is killed.

Two different failure modes:

  • Aug 26 style β€” global slowdown, every operation degraded
  • Aug 31 style β€” normal performance, with a minority of requests hitting something that never returns

If engineering is looking at aggregate latency dashboards, today may well look fine. The p50 is healthy. It is the tail that is killing users, and a p50 chart will hide that completely.

For crews this is arguably worse to work with than uniform slowness. Most saves feel normal, then roughly one in ten just fails, so there is no way to tell in advance whether the record you are entering will survive.


Today’s failures (40 so far, through 14:44 UTC)

Error Count
Data table '<X>' is not accessible… fetching rows from AppSheet database 19
A task was canceled. 8
Unable to find the specified row from AppSheet database 8
Timed out waiting for a resource 3
N rows could not be updated in AppSheet database 2

Nine distinct tables reported unreadable again, ranging from our largest to a static seven-row lookup.

By hour (UTC): 11:00 β†’ 9, 12:00 β†’ 9, 13:00 β†’ 18, 14:00 β†’ 4. Still ongoing at time of posting, and on previous weeks a mid-morning count has roughly tripled by end of day.


Where that leaves the week

Recovery Held for Then
Aug 6 fix 4 working days collapsed to 103 failures on Aug 12
Aug 20 1 day back to 39 on Aug 21
Aug 27–28 2 days back to 40+ today

Three times now a clean stretch has been followed by a bad Monday. Combined with @Jose_Arteaga’s point in #79 and @Rodrigo_Nuvens’ in #67, the simplest reading remains that load exposes this and quiet periods hide it, rather than the fixes landing and then regressing.

@Jose_Arteaga β€” the median/tail split above is the most concrete new thing we have found. If the team is measuring average or p50 sync time, today would look like a good day on our account. It was not. Whatever is happening affects a subset of requests severely rather than all requests mildly, and that seems worth pointing the new telemetry at.

Still happy to share full audit exports, request IDs or stack traces privately if that helps.

Hello all,

First, thank you for getting back to me with your experience. There was a release last Friday and the engineering team was awaiting to get metrics.

@Jonathan_Kitts,

Thank you immensely for this level of detail. Breaking down the p50 median vs. tail latency along with the specific error distributions (especially the resource timeout and unreadable lookup table errors) is exceptionally valuable.

I am sharing this information with the lead engineers investigating the backend connection pooling.

If engineering needs the specific Request IDs or audit exports from that 13:00–14:45 UTC window to trace against the new production telemetry, I will reach out privately.

Thank you again for partnering with us and providing such rigorous data.

Hello all,

Following up with some good news from engineering: another fix has landed, alongside an additional configuration update rolling out today.

Internal metrics are clean and no longer showing the tail-latency timeouts or unreadable table errors.

Since your benchmarks were pivotal in diagnosing this, we’d love to know how your app’s sync look on your end today.

Hello Steve, how do I direct/private message you?

@Jose_Arteaga β€” end-of-day Wednesday. Two full clean days now.

CEILING TIMEOUTS PER DAY  (120s row / 240s sync)

Mon 24 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ             84
Tue 25 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                  72
Wed 26 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  107
Thu 27 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                                      26
Fri 28 Aug  Β·                                                 0
Mon 31 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                      63
Tue 01 Sep  Β·                                                 0   <-- fix lands overnight
Wed 02 Sep  Β·                                                 0


WRITE p90        13:00-14:45 UTC

Fri 28 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                                       25.2s
Mon 31 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ   120.0s  <-- ceiling
Tue 01 Sep  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                                        20.8s
Wed 02 Sep  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                                         20.1s


SYNC p90         13:00-14:45 UTC

Fri 28 Aug  β–ˆ                                                6.1s
Mon 31 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ   215.1s  <-- near 240s ceiling
Tue 01 Sep  β–ˆ                                                6.0s
Wed 02 Sep  β–ˆ                                                6.0s


WRITE VOLUME     13:00-14:45 UTC   (the load being carried)

Fri 28 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                               70
Mon 31 Aug  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                             81
Tue 01 Sep  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ                           88
Wed 02 Sep  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ   184

Wednesday, full working day: 1,300 operations measured, zero exceeded 60
seconds
. Three failures, none of them a timeout β€” all instant rejections with
blank duration, which is our own client double-submitting rather than the
platform. No unreadable-table errors on either day, the first time we have been
clear of those since 12 August.

The last chart is the one I would point engineering at. Wednesday carried
2.3x Monday’s write volume in the same window and produced the best tail
latency of the four days. Monday sat on both ceilings at a third of the load.
Whatever changed, it is not that our traffic dropped.

Worth repeating: the write medians across all four days are 8.8, 7.8, 9.7, 8.8
seconds. Almost flat. The median never moved during this incident, which is why
it was useless as a health signal β€” the tail was the entire symptom.

On timing. Your post said a fix landed Tuesday with a config update rolling
out during that day. Our telemetry has Tuesday clean from the first operation of
the working day (10:29 UTC) straight through the peak, and Monday still failing
at 21:00-22:00 UTC. No gradual transition; it is binary. So the change that
mattered took effect between 31 Aug 22:00 UTC and 1 Sep 10:29 UTC β€” which
suggests the config update still rolling out later on Tuesday is not the one
that fixed it. Might help narrow down which change to credit.

Caveat. Two days. The 6 August fix also gave four clean working days before
12 August came in at 103 ceiling timeouts and 13 August at 226. I am not calling
this resolved until we have a full week at load.

Is everyone else seeing the same?

Thursday 9/3, through 18:00 UTC (2:00 PM ET)

Zero failures. Zero operations over 60 seconds, out of 955. No ceiling timeouts, no unreadable tables, not even an instant rejection.

Third consecutive clean day β€” and the best one yet, at the highest load yet.

13:00–14:45 UTC Fri 8/28 Mon 8/31 Tue 9/1 Wed 9/2 Thu 9/3
Active users 9 10 11 12 11
Writes 70 81 88 184 243
Write median 8.798 s 7.753 s 9.739 s 8.810 s 6.924 s
Write p90 25.151 s 120.006 s 20.798 s 20.074 s 16.409 s
Write max 32.663 s 120.180 s 29.691 s 28.586 s 25.198 s
Writes > 60 s 0 14 0 0 0
Sync p90 6.061 s 215.077 s 5.950 s 6.042 s 5.274 s
Sync max 7.497 s 240.583 s 8.172 s 10.172 s 9.399 s

So far, we haven’t seen any new errors, and it looks like we’ve seen a significant improvement. Now, we just need to wait until Monday, when we reach the peak usage period, to confirm whether the issue has been fully resolved.

We’ll continue monitoring things closely on our end. Thank you very much for your support and assistance!

it seems that definitely this problem has been solved thanks @Jose_Arteaga

Thank you Jose,

I appreciate you for being patient and also for being proactive!