2 months ago
i’m sending a json POST to an api at data.lacity.org to pull some data. this works fine 100% of the time when i run the app locally. however when i run it from railway, depending on the deploy, i get 403 errors:
7/20/2026, 6:31:41 PM: api: /api/v3/views/73a2-6ar5/query.json postData: {"query":"SELECT locator_gis_returned_address, casenumber, status, createddate, systemmodstamp, closeddate, resolution_code__c, reason_code__c, type, geolocation__latitude__s, geolocation__longitude__s WHERE type = 'Streetlight Repair Services' and systemmodstamp > '2025-04-01' and geolocation__latitude__s > 34.062125 and geolocation__latitude__s < 34.083554 and geolocation__longitude__s > -118.338907 and geolocation__longitude__s < 118.326518","includeSynthetic":false}
7/20/2026, 6:31:42 PM: postJson error: 403 "\r\n403 Forbidden\r\n\r\n403 Forbidden\r\nnginx\r\n\r\n\r\n"
I’m using axios to post the query. The apiToken doesn’t change.
const axionApi = axios.create({
baseURL: 'https://data.lacity.org',
responseType: 'json',
headers: {
'Content-Type' : 'application/json',
'X-App-Token' : apiToken,
'User-Agent': 'Mozilla/5.0 (Node.js LA City Query)'
}});
The explanation I’ve gotten so far is that data.lacity.org are using a WAF on nginx, and some of the railway URLs are blacklisted. I have no way to contact them about this and no idea why this would be the case. Can you suggest a solution?
6 Replies
2 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 3 months ago
2 months ago
Your 403 body is nginx's default HTML error page. That's the detail worth pulling on — Socrata returns JSON errors for anything auth-related, including a bad, missing, or rate-limited X-App-Token. A bare nginx HTML 403 means the request was rejected at the edge, before it ever reached the Socrata application.
So your app token isn't the problem and rotating it won't help. This is IP reputation, which also explains the pattern you noticed: it varies by deploy because your outbound address varies by deploy. Some of the addresses you land on are on a blocklist, some aren't. Your instinct about the WAF was right.
Two things worth doing.
First, confirm it cheaply. Add a line at startup that hits any IP-echo endpoint and logs the result, then correlate the egress address against which deploys fail. If the failures cluster on particular addresses, that's your confirmation — and it's also exactly what you'd send to whoever can unblock it.
Second, try the classic SODA endpoint rather than the v3 query API:
https://data.lacity.org/resource/73a2-6ar5.json?$query=
I just tested that dataset and it returns 200 with no app token at all. Different path, potentially different WAF rules, and it takes ten minutes to find out whether the block is host-wide or specific to /api/v3/views/. Your SoQL carries over unchanged, so it's a cheap test.
If both paths are blocked then it's purely IP-level, and your realistic options are to route that one outbound call through a proxy on an address that isn't blocked, or to get the range allowlisted.
On contact — you do have a route. data.lacity.org is Socrata-hosted (now Tyler Technologies) and the DataLA portal has a support/feedback channel. Worth one message describing the block with the specific addresses, because a WAF rule that catches an entire cloud provider's range is almost always unintentional collateral rather than a deliberate policy aimed at you.
2 months ago
i did add code to print the external IP, but i’ve redeployed 3 times and it is now not changing from 162.220.232.72, which works fine...
I will try the SODA query to see if that makes any difference.
2 months ago
That's a useful result, and it corrects something I said. I claimed your outbound address varies per deploy — three redeploys holding at 162.220.232.72 says it doesn't, at least not reliably. Scratch that part.
Which actually makes the problem more tractable, because it means the variable isn't which IP you land on. Same address, works now, 403'd before. So it's time-varying, not deploy-varying — a WAF rule or blocklist that 162.220.232.72 drifts on and off of, or something request-shaped rather than origin-shaped.
There's a sharper test available. Socrata's real responses carry identifying headers — I checked, and both a 200 from the classic endpoint and a 200 from the v3 query endpoint come back with:
Server: nginx
X-Socrata-Region: aws-us-east-1-fedramp-prod
X-Socrata-RequestId:
So log the full response headers on your next 403, and look for X-Socrata-RequestId.
Present on the 403 → the request reached Socrata and Socrata refused it. That's app-level: token rejected, quota, or a per-account rule. Different problem entirely, and one you can raise with a request ID in hand, which is the thing that makes support tickets go fast.
Absent on the 403 → it died at the edge before Socrata ever saw it. That's the WAF theory confirmed, and the request ID's absence is your evidence.
That single header splits it cleanly and costs you one console.log.
While you're at it, worth confirming the token is actually being honored rather than silently ignored — an invalid X-App-Token doesn't error, it just drops you to anonymous limits, which would make bursts fail intermittently while single requests look fine.
I also ran 15 rapid requests against that dataset from a clean address and got 200 on every one, so there's no aggressive burst throttle on the endpoint itself. Whatever's rejecting you is more specific than simple rate limiting.
Curious what the SODA endpoint does for you — if it works while the v3 path 403s from the same IP at the same moment, that's the strongest possible signal that the rule is path-scoped rather than address-scoped, and you can just move to the classic endpoint and be done.
myiephero
That's a useful result, and it corrects something I said. I claimed your outbound address varies per deploy — three redeploys holding at 162.220.232.72 says it doesn't, at least not reliably. Scratch that part. Which actually makes the problem more tractable, because it means the variable isn't which IP you land on. Same address, works now, 403'd before. So it's time-varying, not deploy-varying — a WAF rule or blocklist that 162.220.232.72 drifts on and off of, or something request-shaped rather than origin-shaped. There's a sharper test available. Socrata's real responses carry identifying headers — I checked, and both a 200 from the classic endpoint and a 200 from the v3 query endpoint come back with: Server: nginx X-Socrata-Region: aws-us-east-1-fedramp-prod X-Socrata-RequestId: <hex> So log the full response headers on your next 403, and look for X-Socrata-RequestId. Present on the 403 → the request reached Socrata and Socrata refused it. That's app-level: token rejected, quota, or a per-account rule. Different problem entirely, and one you can raise with a request ID in hand, which is the thing that makes support tickets go fast. Absent on the 403 → it died at the edge before Socrata ever saw it. That's the WAF theory confirmed, and the request ID's absence is your evidence. That single header splits it cleanly and costs you one console.log. While you're at it, worth confirming the token is actually being honored rather than silently ignored — an invalid X-App-Token doesn't error, it just drops you to anonymous limits, which would make bursts fail intermittently while single requests look fine. I also ran 15 rapid requests against that dataset from a clean address and got 200 on every one, so there's no aggressive burst throttle on the endpoint itself. Whatever's rejecting you is more specific than simple rate limiting. Curious what the SODA endpoint does for you — if it works while the v3 path 403s from the same IP at the same moment, that's the strongest possible signal that the rule is path-scoped rather than address-scoped, and you can just move to the classic endpoint and be done.
2 months ago
actually i am not sure what the external IP was when it failed - it has not failed since I added the code to show the external IP. I will redeploy a few more times to see if i can get it to change and fail.
2 months ago
That's still a useful data point, even though it's a null one.
Worth noticing what you've already established though: the address held at 162.220.232.72 across three redeploys. Redeploying is the one lever you now have evidence doesn't move the variable, so you're unlikely to reproduce it that way and you could burn a lot of deploys finding that out. I'd stop hunting and make the capture passive instead.
Concretely, move the logging out of startup and into the axios error path so it fires whenever the next 403 lands rather than only when you happen to be watching:
catch (err) {
if (err.response) {
console.log(new Date().toISOString(), err.response.status, JSON.stringify(err.response.headers), String(err.response.data).slice(0, 300));}
}
The header dump is the part that matters. X-Socrata-RequestId present means Socrata saw the request and refused it; absent means it died at the edge. I re-ran both endpoints from a clean address just now and both still return 200 with no app token and both still carry that header, so the test is valid and the v3 path isn't globally blocked.
Also worth saying plainly: your last observed failure was 20 July. Nine clean days is long enough that this may already have been fixed upstream without anyone telling you. Instrument and wait beats trying to force it.
One unrelated thing I noticed, since the query is in your post:
geolocation__longitude__s > -118.338907 and geolocation__longitude__s < 118.326518
The upper bound is missing its minus sign, so instead of a narrow strip you're spanning everything from -118.34 eastward round to +118.33. I ran both against that dataset: as written it matches 3164 rows, with the minus restored it matches 182. That won't cause an nginx 403 and I'm not suggesting it's related, but you're pulling about seventeen times the data you intend on every call.
2 months ago
i’ve added the header dump to the catch block as suggested.
not sure where you were seeing the incorrect longitude - it might have been a typo in my test script that I posted here and then deleted. the actual posts look like this (from the production logs):
7/28/2026, 9:35:46 PM: api: /api/v3/views/73a2-6ar5/query.json postData: {"query":"SELECT locator_gis_returned_address, casenumber, status, createddate, systemmodstamp, closeddate, resolution_code__c, reason_code__c, type, geolocation__latitude__s, geolocation__longitude__s WHERE type = 'Streetlight Repair Services' and systemmodstamp > '2025-04-01' and geolocation__latitude__s > 34.062125 and geolocation__latitude__s < 34.083554 and geolocation__longitude__s > -118.338907 and geolocation__longitude__s < -118.326518","includeSynthetic":false}
7/28/2026, 9:35:48 PM: api: /api/v3/views/2cy6-i7zn/query.json postData: {"query":"SELECT locator_gis_returned_address, casenumber, status, createddate, systemmodstamp, closeddate, resolution_code__c, reason_code__c, type, geolocation__latitude__s, geolocation__longitude__s WHERE type = 'Streetlight Repair Services' and systemmodstamp > '2025-04-01' and geolocation__latitude__s > 34.062125 and geolocation__latitude__s < 34.083554 and geolocation__longitude__s > -118.338907 and geolocation__longitude__s < -118.326518","includeSynthetic":false}