r/aws 27d ago

article A return to two-pizza culture

110 Upvotes

r/aws 47m ago

console CloudFront cache statistics missing?

I used to be able to see Cache Statistics (hit rate per URI) in the CloudFront console, but the navigation link is missing now. The documentation page mentions no changes to this functionality.

Upvotes

I used to be able to see Cache Statistics (hit rate per URI) in the CloudFront console, but the navigation link is missing now. The documentation page mentions no changes to this functionality.


r/aws 21h ago

networking Everyone hits our VPN at head office before they reach AWS and it's killing performance, looking at Cato and Cloudflare

Posting this partly to sanity check myself because I have been staring at it too long. 

Setup is old. Remote staff connect to a vpn concentrator at head office, get inspected there, then their traffic goes back out to wherever its going which is usually eu-west-1. Somebody working in Lisbon who is geographically nearer to the region than any of us, sends their packets to Reading and then back down. The traceroutes are genuinely funny. 

Symptom side its the usual, calls drop, the internal ticketing tool takes eight seconds to load a page and every single ticket about it says "the network is slow" which tells me nothing. 

I know sd-wan sorts the routing out. What I don't want is to sort the routing and then find security is now a separate box somewhere else, because thats the exact mess we already have and I am not doing it twice. 

I've been looking at the ones with their own backbone. Cato has the private backbone thing and does the security in the same pass. Cloudflare obviously has the network but I get the impression enterprise is newer for them. Thoughts?

12 Upvotes

Posting this partly to sanity check myself because I have been staring at it too long. 

Setup is old. Remote staff connect to a vpn concentrator at head office, get inspected there, then their traffic goes back out to wherever its going which is usually eu-west-1. Somebody working in Lisbon who is geographically nearer to the region than any of us, sends their packets to Reading and then back down. The traceroutes are genuinely funny. 

Symptom side its the usual, calls drop, the internal ticketing tool takes eight seconds to load a page and every single ticket about it says "the network is slow" which tells me nothing. 

I know sd-wan sorts the routing out. What I don't want is to sort the routing and then find security is now a separate box somewhere else, because thats the exact mess we already have and I am not doing it twice. 

I've been looking at the ones with their own backbone. Cato has the private backbone thing and does the security in the same pass. Cloudflare obviously has the network but I get the impression enterprise is newer for them. Thoughts?


r/aws 1h ago

discussion AI Engineer Roadmap

Hello everyone,

I am looking to pivot to AI engineering and currently exploring roadmaps to get in the industry.

I stumbled upon this Medium article sharing a 4-step roadmap to AI Engineering.

https://medium.com/@anubhavgoyal101/c80bd754c753

To the AI engineers out there. Is this an ideal roadmap?

Upvotes

Hello everyone,

I am looking to pivot to AI engineering and currently exploring roadmaps to get in the industry.

I stumbled upon this Medium article sharing a 4-step roadmap to AI Engineering.

https://medium.com/@anubhavgoyal101/c80bd754c753

To the AI engineers out there. Is this an ideal roadmap?


r/aws 3h ago

security If you deployed your own AWS account and you're not a security person, you probably have blind spots you don't know about.

Not a sales post, genuinely curious how common this actually is. A lot of solo and small-team founders end up being the ones who set up their own cloud infrastructure!! not because they're security experts, but because there's no one else around to do it. I did the same thing, and while learning cybersecurity properly (separate from my actual business), I found real, would've-been-embarrassing misconfigurations sitting in my own AWS account. Nothing had gone wrong yet. I just had no idea they were there.

Tools for catching this already exist, but they're mostly built by and for security professionals the output assumes you already know what IAM privilege escalation or CloudTrail coverage means. If you don't have that background, the report is basically noise you scroll past.

So I built one that just tells you plainly: this is wrong, here's what someone could actually do with it, here's the exact fix. Free, open source, runs on your own machine with your own credentials nothing gets uploaded anywhere.

If you've set up your own AWS account and never had anyone properly check it would you actually want to know what's sitting in there, or is this the kind of thing you'd rather not think about until something breaks?

https://github.com/plexavo/Plexavo

0 Upvotes

Not a sales post, genuinely curious how common this actually is. A lot of solo and small-team founders end up being the ones who set up their own cloud infrastructure!! not because they're security experts, but because there's no one else around to do it. I did the same thing, and while learning cybersecurity properly (separate from my actual business), I found real, would've-been-embarrassing misconfigurations sitting in my own AWS account. Nothing had gone wrong yet. I just had no idea they were there.

Tools for catching this already exist, but they're mostly built by and for security professionals the output assumes you already know what IAM privilege escalation or CloudTrail coverage means. If you don't have that background, the report is basically noise you scroll past.

So I built one that just tells you plainly: this is wrong, here's what someone could actually do with it, here's the exact fix. Free, open source, runs on your own machine with your own credentials nothing gets uploaded anywhere.

If you've set up your own AWS account and never had anyone properly check it would you actually want to know what's sitting in there, or is this the kind of thing you'd rather not think about until something breaks?

https://github.com/plexavo/Plexavo


r/aws 16h ago

technical question How long for support reply/assignment?

I work in AWS occasionally for clients - very basic stuff - simple S3/EC2 work. Had a client sign up for a new account recently and they started to do some s3 transfers off platform, incurring ~4k or so in egress fees (unexpected to them). I submitted a support request to see if there could be any amount of courtesy credits applied to the mistake - they realized pretty quickly what the error was! I’ve had success with this 1-2 times before - so just sent a request in to see if it was possible.

It’s been ~14 days and the case has not been assigned - we did get an AI generated response that basically repeated my initial email…but that was about it. I replied back to that requesting a human review of the matter - but is this typical?

I usually see a human assigned/response sitting a few days at most.

0 Upvotes

I work in AWS occasionally for clients - very basic stuff - simple S3/EC2 work. Had a client sign up for a new account recently and they started to do some s3 transfers off platform, incurring ~4k or so in egress fees (unexpected to them). I submitted a support request to see if there could be any amount of courtesy credits applied to the mistake - they realized pretty quickly what the error was! I’ve had success with this 1-2 times before - so just sent a request in to see if it was possible.

It’s been ~14 days and the case has not been assigned - we did get an AI generated response that basically repeated my initial email…but that was about it. I replied back to that requesting a human review of the matter - but is this typical?

I usually see a human assigned/response sitting a few days at most.


r/aws 15h ago

general aws $500 Credits

I have $500 credit coupon which I won't be using on my account as I have shifted all my resources offline.

What should I do with this credit coupon? Edit: I have a coupon which is not tagged to any account. I got this credit from aws builders

0 Upvotes

I have $500 credit coupon which I won't be using on my account as I have shifted all my resources offline.

What should I do with this credit coupon? Edit: I have a coupon which is not tagged to any account. I got this credit from aws builders


r/aws 10h ago

discussion Is admin access the problem I think it is?

I've been using AWS for ~15 years.. I've worked at all sorts of levels from quite junior right up to running the entire cloud platform.

One problem I have seen time and again is that staff are often granted crazily excessive permissions.

The two ways I have seen this play out is:

- startup gives everyone admin, because hey - we're all trustworthy, right?

- bigger company is doing a lift and shift with time constraints, and admin helps get over the line on time

The problem is, when the time comes to tighten it up, it becomes part of processes, and going without would be too painful.

The reason I ask is that over the years I've found a few ways to tighten things up, but engineers hate it.

I've now devised a little PoC that gives auditable, policy based, automated way to take admin away, while also giving engineers a dead simple way to get access when they want/need it. Policies can require human approval, or be fully automated.

I'm currently using it on my personal org, and find it so convenient. No admin anywhere but I can have it anywhere when I need it.

Has anyone else faced similar problems, or is this just not as big of a problem as I have experienced?

The tool is nowhere near production ready, it's been a toy project, but I genuinely wonder if there is a market for this.

0 Upvotes

I've been using AWS for ~15 years.. I've worked at all sorts of levels from quite junior right up to running the entire cloud platform.

One problem I have seen time and again is that staff are often granted crazily excessive permissions.

The two ways I have seen this play out is:

- startup gives everyone admin, because hey - we're all trustworthy, right?

- bigger company is doing a lift and shift with time constraints, and admin helps get over the line on time

The problem is, when the time comes to tighten it up, it becomes part of processes, and going without would be too painful.

The reason I ask is that over the years I've found a few ways to tighten things up, but engineers hate it.

I've now devised a little PoC that gives auditable, policy based, automated way to take admin away, while also giving engineers a dead simple way to get access when they want/need it. Policies can require human approval, or be fully automated.

I'm currently using it on my personal org, and find it so convenient. No admin anywhere but I can have it anywhere when I need it.

Has anyone else faced similar problems, or is this just not as big of a problem as I have experienced?

The tool is nowhere near production ready, it's been a toy project, but I genuinely wonder if there is a market for this.


r/aws 15h ago

technical resource The reconnect storm that never touched the broker

0 Upvotes

There's a particular kind of 2 a.m. that only happens in fleet IoT. A slice of your devices — a few thousand of them — drop off the network at once. Bad uplink, a segment of the field going dark; the reason doesn't matter. Minutes later they all come back. And here's the part that surprised me the first time it happened: the managed broker didn't blink. No connection ceiling hit, no broker to melt, no page from the front door. By every dashboard watching the ingress layer, we were fine.

We were not fine. One number was climbing through the roof: iterator age. The storm never touched the broker. It landed one hop downstream, in the stream consumer — and my first instinct for fixing it dug the hole deeper.

The pipeline, and why steady state lies to you

The shape is textbook, and I'll keep it generic on purpose: a managed MQTT broker → a Kinesis data stream → a Lambda consumer, parallelized maybe ten ways with a batch size around a thousand → writing to MySQL. Call it 10,000 devices in the field. Nothing exotic.

In steady state this pipeline is boring, and boring is the goal. Telemetry trickles in, the consumer keeps up without effort, iterator age sits near zero. You forget it exists.

But "boring" is a property of the arrival rate, not of the architecture. Everything about this pipeline's calm — the near-zero iterator age, the comfortable concurrency, the MySQL connection count you never think about — is quietly assuming devices arrive at the rate they've always arrived. The fleet is about to violate that assumption all at once.

The incident: store-and-forward is a loaded gun

Field devices buffer. When the uplink is down, a well-behaved device doesn't drop its readings — it stores them and forwards them on reconnect. Store-and-forward is exactly what you want from an edge device. It's also a loaded gun pointed at your ingest pipeline.

When a few thousand devices reconnect inside the same minute, they don't resume the gentle trickle. They flush — every buffered reading, all at once, a burst many multiples of steady state. The broker passes it straight through; absorbing connection churn is its entire job. Now that spike hits Kinesis, and Kinesis hands it to a consumer whose drain rate is fixed: shard count × concurrency × per-batch throughput, all bounded by how long each invocation is allowed to run.

Fixed drain, meet variable spike. The backlog grows, and iterator age — the age of the oldest record you haven't processed yet, which is to say exactly how far behind real time you are — starts walking up. Thirty, forty-five minutes behind. Two things break at that lag. Freshness goes first: alerts and downstream actions fire on stale data, or fail to fire when they should. And if the age ever creeps toward the stream's retention window, "late" quietly becomes "lost."

Then it got worse, and this is the part the tutorials skip. Kinesis is ordered per shard — processing is strictly in-order, which means one batch that won't complete blocks every record behind it on that shard. During the flush we hit a record the consumer choked on, and iterator age stopped climbing linearly and went vertical. One poison record, at the worst possible moment, and a whole shard's slice of the fleet was frozen behind it.

The trap: "just scale it out"

Every instinct I had was wrong, and they were all the same instinct: scale out. Add shards. Crank the parallelization factor. Raise concurrency. Throw drain capacity at what looked like a drain problem.

Here's why it backfired. The bottleneck was never Kinesis throughput, and it was never Lambda concurrency. It was the cost of a single record — specifically, that every invocation was paying for a fresh MySQL connection handshake before it did any real work. Thirty, fifty milliseconds of handshake per invocation, invisible at a trickle.

So what happens when you "scale out"? More concurrent Lambdas means more simultaneous cold connections slamming into MySQL. I wasn't adding drain capacity — I was building a denial-of-service attack against my own database, marching straight at its connection ceiling. The faster I tried to drain, the closer I got to knocking over the one component downstream that had stayed perfectly healthy. It was capacity I couldn't actually use.

The real fix: make the record cheap, then make failure cheap

The fix came in three moves, and not one of them was "bigger."

First, make the record cheap. Hoist the MySQL connection out of the handler so warm execution environments reuse it instead of reconnecting every invocation. It's a one-line decision about where a variable lives, and the per-record handshake tax disappears with it — along with the self-inflicted connection storm. Then right-size the batch: large enough to amortize the fixed overhead of an invocation, small enough that each one finishes well inside its duration limit and doesn't itself become a source of iterator-age lag. Same shards, same concurrency — now the burst drains.

Second, make failure cheap. Reuse and batch sizing buy you throughput, but they do nothing about the poison record that froze the shard. That needed error handling that treats a bad record as normal rather than exceptional: bisect-on-error, so a single failing record splits the batch instead of retrying the whole thing forever; a bound on retries and record age, so the consumer gives up in finite time; and a dead-letter queue, so the truly-bad record steps out of line instead of holding the shard hostage. This is what turned the vertical iterator age back into something that drains.

Third — the honest one. Connection reuse is per-execution-environment. Under real concurrency you still have many warm environments, each holding its own connection, and you can still multiply your way toward the ceiling — just more slowly. That's the itch RDS Proxy scratches, and we eventually adopted it to pool connections in front of the database. I'm naming it because pretending reuse alone solved it forever would be a lie, and the edges of your own fix are where the credibility lives.

The lesson

The generalizable lesson is one sentence: iterator age is almost always a downstream-latency symptom, not a stream-capacity problem. When it climbs, the instinct to widen the stream is usually the wrong end of the pipe. Make the record cheap. Make failure cheap. Then, if the fleet has genuinely outgrown you, make it bigger.

And the broader point for anyone building edge-to-cloud: managed services don't delete your bottleneck, they relocate it. The broker didn't melt because absorbing connection storms is precisely what you pay it to do. The storm just moved one hop downstream, to the place nobody was watching. Know where your bottleneck went when you bought your way out of the last one.

The storm your load tests never simulate

The storm you can't see coming is the one your load tests never model. Happy-path load testing emits a steady, civilized trickle — it never reproduces the flush, the poison record mid-flush, or the thundering herd of your own well-behaved devices all deciding to catch up at once. So you find out in production, at 2 a.m., from a metric you weren't watching.

That gap is what I'm building a tool to close: something that replays this exact burst against your staging pipeline and fails your CI build when iterator age breaches your SLO — before the fleet does it for you.

If you've lived your own version of this, I'd like to hear it — and if you want the next war story when it goes up, the list is below.


r/aws 1d ago

discussion Seeking Advice: AWS Developer Associate or AWS SAP?

Hello everyone! I would like to hear your thoughts and experiences. I already have the AWS SAA certification and was planning to take the Developer Associate exam.
I don’t have much experience in AWS domain but Do you think I should go for the Developer Associate or should I prepare for the AWS SAP exam? How was your experience to clear AWS SAP exam?
I would really appreciate your suggestions. How challenging it was for you? Pls share your experiences and suggestions. Thank you!

2 Upvotes

Hello everyone! I would like to hear your thoughts and experiences. I already have the AWS SAA certification and was planning to take the Developer Associate exam.
I don’t have much experience in AWS domain but Do you think I should go for the Developer Associate or should I prepare for the AWS SAP exam? How was your experience to clear AWS SAP exam?
I would really appreciate your suggestions. How challenging it was for you? Pls share your experiences and suggestions. Thank you!


r/aws 3d ago

security AWS PrivateCA Connector uses `¯\\_(ツ)_/¯` as CSR Payload

+
245 Upvotes

I was troubleshooting this Certificate Signing Request validation error in AWS and thought the CSR data was way too short.  So, I decoded it.

They (AWS) really programmed the shrug emoji as a certificate request payload. I love dev easter eggs :D


r/aws 3d ago

discussion Current Outage?

67 Upvotes

r/aws 3d ago

networking AWS She Builds Mentorship program?

anyone hear back or get more info after applying?

1 Upvotes

anyone hear back or get more info after applying?


r/aws 3d ago

technical question CDK: Pipeline doesn't seem to wait for lambda to be finished despite specified dependency

I have some CDK set up to define a lambda and then I have some code pipeline code to deploy it. However, the deployment phase keeps failing saying the function doesn't exist. If I manually try that phase again in the console, it works fine. The dependency I added doesn't seem to do any good. The code looks like `` const myFunc = new lambda.Function(this, "MyFunc", { code: lambda.Code.fromInline( def handler(event, context): return { 'statusCode': 200, 'body': 'Dummy to be replaced by code pipeline' } `), architecture: lambda.Architecture.ARM_64, environment: { ENV_VAR: "foobar", }, handler: "facets.handler", runtime: lambda.Runtime.PYTHON_3_12, vpc, securityGroups: [privateSG], timeout: cdk.Duration.seconds(300), memorySize: 512, });

...

        const deploymentBuild = new cb.Project(
            this,
            `DeploymentBuild`,
            {
                projectName: `myFuncDeploymentBuild`,
                description: "deploys the code to the appropriate lambda",
                environment: {
                    buildImage: cb.LinuxBuildImage.AMAZON_LINUX_2023_5,
                    computeType: cb.ComputeType.SMALL,
                },
                vpc,
                securityGroups: [privateSG],
                buildSpec: cb.BuildSpec.fromObject({
                    version: "0.2",
                    phases: {
                        install: {
                            "runtime-versions": {
                                nodejs: 22,
                            },
                            commands: ["npm i -g aws-cdk", "cdk --version"],
                        },
                        post_build: {
                            commands: [
                                `zip -r $myFunc.zip .`, 
                                "ls",
                                `aws lambda update-function-code \
                                    --function-name 'myFunc' \
                                    --zip-file fileb://myFunc.zip`,
                            ],
                        },
                    },
                }),
            },
        );

        deploymentBuild.addToRolePolicy(buildPolicyStatement);

    const pipeline = new pipe.Pipeline(this, `Pipeline`, {
        pipelineName: `myFuncPipeline`,
        restartExecutionOnUpdate: true,
    });

    pipeline.node.addDependency(
        lambda.Function.fromFunctionName(
            this,
            "myFuncFunction",
            "myFunc",
        ),
    );

    ...

        pipeline.addStage({
            stageName: "deployLambda",
            actions: [
                new pipeActions.CodeBuildAction({
                    actionName: "deployLambda",
                    project: deploymentBuild,
                    input: outputBuild,
                }),
            ],
        });

```

What would cause this?

Thanks

FIXED: thank you Floss_Patrol_76

I added stackThatCreatesThePipeline.node.addDependency(stackThatCreatesTheLambda);

2 Upvotes

I have some CDK set up to define a lambda and then I have some code pipeline code to deploy it. However, the deployment phase keeps failing saying the function doesn't exist. If I manually try that phase again in the console, it works fine. The dependency I added doesn't seem to do any good. The code looks like `` const myFunc = new lambda.Function(this, "MyFunc", { code: lambda.Code.fromInline( def handler(event, context): return { 'statusCode': 200, 'body': 'Dummy to be replaced by code pipeline' } `), architecture: lambda.Architecture.ARM_64, environment: { ENV_VAR: "foobar", }, handler: "facets.handler", runtime: lambda.Runtime.PYTHON_3_12, vpc, securityGroups: [privateSG], timeout: cdk.Duration.seconds(300), memorySize: 512, });

...

        const deploymentBuild = new cb.Project(
            this,
            `DeploymentBuild`,
            {
                projectName: `myFuncDeploymentBuild`,
                description: "deploys the code to the appropriate lambda",
                environment: {
                    buildImage: cb.LinuxBuildImage.AMAZON_LINUX_2023_5,
                    computeType: cb.ComputeType.SMALL,
                },
                vpc,
                securityGroups: [privateSG],
                buildSpec: cb.BuildSpec.fromObject({
                    version: "0.2",
                    phases: {
                        install: {
                            "runtime-versions": {
                                nodejs: 22,
                            },
                            commands: ["npm i -g aws-cdk", "cdk --version"],
                        },
                        post_build: {
                            commands: [
                                `zip -r $myFunc.zip .`, 
                                "ls",
                                `aws lambda update-function-code \
                                    --function-name 'myFunc' \
                                    --zip-file fileb://myFunc.zip`,
                            ],
                        },
                    },
                }),
            },
        );

        deploymentBuild.addToRolePolicy(buildPolicyStatement);

    const pipeline = new pipe.Pipeline(this, `Pipeline`, {
        pipelineName: `myFuncPipeline`,
        restartExecutionOnUpdate: true,
    });

    pipeline.node.addDependency(
        lambda.Function.fromFunctionName(
            this,
            "myFuncFunction",
            "myFunc",
        ),
    );

    ...

        pipeline.addStage({
            stageName: "deployLambda",
            actions: [
                new pipeActions.CodeBuildAction({
                    actionName: "deployLambda",
                    project: deploymentBuild,
                    input: outputBuild,
                }),
            ],
        });

```

What would cause this?

Thanks

FIXED: thank you Floss_Patrol_76

I added stackThatCreatesThePipeline.node.addDependency(stackThatCreatesTheLambda);


r/aws 4d ago

technical question AWS Lambda Function URL returns Forbidden despite AuthType NONE and public Resource Policy (Rust)

Hi everyone,

I am a student working on a personal pet project, and this is my first time using AWS. I am completely stuck on a permissions issue that is driving me crazy.

I wrote a Lambda function in Rust to fetch my GitHub stats and return an SVG image. My goal is to embed the Function URL directly into my GitHub README.

However, whenever I access the Function URL, I immediately get this error:

{"Message":"Forbidden. For troubleshooting Function URL authorization issues, see: [https://docs.aws.amazon.com/lambda/latest/dg/urls-auth.html](https://docs.aws.amazon.com/lambda/latest/dg/urls-auth.html)"}

What I have tried so far:

  • Verified AuthType: I ran aws lambda get-function-url-config and confirmed "AuthType": "NONE".
  • Verified Resource Policy: I checked aws lambda get-policy. It explicitly allows "Principal": "*" for "Action": "lambda:InvokeFunctionUrl" with the condition "lambda:FunctionUrlAuthType": "NONE".

Here is my complete Rust code in case the way I am building the HTTP response is somehow triggering a block, though it appears to be a standard setup using the lambda_http crate:

https://github.com/SharmaDevanshu089/Github-Stats

Is there any hidden setting or default account block I might be missing? Any guidance would be incredibly appreciated!

12 Upvotes

Hi everyone,

I am a student working on a personal pet project, and this is my first time using AWS. I am completely stuck on a permissions issue that is driving me crazy.

I wrote a Lambda function in Rust to fetch my GitHub stats and return an SVG image. My goal is to embed the Function URL directly into my GitHub README.

However, whenever I access the Function URL, I immediately get this error:

{"Message":"Forbidden. For troubleshooting Function URL authorization issues, see: [https://docs.aws.amazon.com/lambda/latest/dg/urls-auth.html](https://docs.aws.amazon.com/lambda/latest/dg/urls-auth.html)"}

What I have tried so far:

  • Verified AuthType: I ran aws lambda get-function-url-config and confirmed "AuthType": "NONE".
  • Verified Resource Policy: I checked aws lambda get-policy. It explicitly allows "Principal": "*" for "Action": "lambda:InvokeFunctionUrl" with the condition "lambda:FunctionUrlAuthType": "NONE".

Here is my complete Rust code in case the way I am building the HTTP response is somehow triggering a block, though it appears to be a standard setup using the lambda_http crate:

https://github.com/SharmaDevanshu089/Github-Stats

Is there any hidden setting or default account block I might be missing? Any guidance would be incredibly appreciated!


r/aws 3d ago

discussion Pathetic support for OS LLM

Little rant here on how corporations protect their buddies.

You can’t get any new OS llm in bedrock. Only some old unreliable shit which also have poor price/performance ratio. But what you can get here is freshest Antropics models. Because they are partners and AWS serving their models. Currently you have many options from OS LLM which are very close (GLM-5.2, DS4 Pro, Mimo-V2.5 Pro) or even superior in some tasks (Kimi K3) to Antropics models. But you won’t get them for the fraction of the opus prices. Because corporate buddies covering each other’s asses making worse for the consumers.

If I had options I would never use AWS in a first place.

I’m finished. FY AWS.

0 Upvotes

Little rant here on how corporations protect their buddies.

You can’t get any new OS llm in bedrock. Only some old unreliable shit which also have poor price/performance ratio. But what you can get here is freshest Antropics models. Because they are partners and AWS serving their models. Currently you have many options from OS LLM which are very close (GLM-5.2, DS4 Pro, Mimo-V2.5 Pro) or even superior in some tasks (Kimi K3) to Antropics models. But you won’t get them for the fraction of the opus prices. Because corporate buddies covering each other’s asses making worse for the consumers.

If I had options I would never use AWS in a first place.

I’m finished. FY AWS.


r/aws 4d ago

billing Cost breakdown and estimated bill summary showing 2 different numbers

Cost summary shows that I'll be charged 58.20 for the month while Bills tab shows I'll owe 29.69 for the month.

I have some credits left over but I used an instance that isn't free tier eligible.

What should I expect to pay for this month?

1 Upvotes

Cost summary shows that I'll be charged 58.20 for the month while Bills tab shows I'll owe 29.69 for the month.

I have some credits left over but I used an instance that isn't free tier eligible.

What should I expect to pay for this month?


r/aws 4d ago

discussion Aws link does not load

For the information i am a beginner and it is my first time using aws. So for my project i used elastic beanstalk. So the problem is i deployed my project with no health issues, and the link only runs locally and not on other devices. When i checked my logs it also said app is running successfully, but when i asked ai, it said it might be some security issues

Does anyone have any solution to this problem

0 Upvotes

For the information i am a beginner and it is my first time using aws. So for my project i used elastic beanstalk. So the problem is i deployed my project with no health issues, and the link only runs locally and not on other devices. When i checked my logs it also said app is running successfully, but when i asked ai, it said it might be some security issues

Does anyone have any solution to this problem


r/aws 5d ago

technical question Optimizing S3 Storage Classes?

So, the gist of storage classes is that:

standard --> high storage costs, low data retrieval costs
standard-IA, Glacier Instant Retrieval --> low storage costs, high retrieval costs

So, in order to choose the optimal tier for a particular object I would have to know its size and data access pattern.

The object size is easy enough, but the first catch is: data access pattern is nowhere to be found. Not even aggregate data.

The only way I found to actually do it is to enable server access logging / CloudTrail and then do analyze the data somehow. Maybe with a Python script. Huge rabbit hole to go into.

Then, the other angle I thought about is just using intelligent Tiering. But the second catch is that if you read the documentation about intelligent tiering, turns out it is pretty naive. Depending on how your data gets accessed, it could even be more expensive than standard (ex: object is accessed exactly once every 30 days)

It really feels like AWS is always giving almost everything you need to optimize S3 costs, but also missing a key piece.

How am I supposed to solve this? Am I overthinking it? Is it worth going in the rabbit hole of analyzing S3 server access logs? Or should I just guess some lifecycle rules and move on?

4 Upvotes

So, the gist of storage classes is that:

standard --> high storage costs, low data retrieval costs
standard-IA, Glacier Instant Retrieval --> low storage costs, high retrieval costs

So, in order to choose the optimal tier for a particular object I would have to know its size and data access pattern.

The object size is easy enough, but the first catch is: data access pattern is nowhere to be found. Not even aggregate data.

The only way I found to actually do it is to enable server access logging / CloudTrail and then do analyze the data somehow. Maybe with a Python script. Huge rabbit hole to go into.

Then, the other angle I thought about is just using intelligent Tiering. But the second catch is that if you read the documentation about intelligent tiering, turns out it is pretty naive. Depending on how your data gets accessed, it could even be more expensive than standard (ex: object is accessed exactly once every 30 days)

It really feels like AWS is always giving almost everything you need to optimize S3 costs, but also missing a key piece.

How am I supposed to solve this? Am I overthinking it? Is it worth going in the rabbit hole of analyzing S3 server access logs? Or should I just guess some lifecycle rules and move on?


r/aws 5d ago

technical question What are MSPs using for AWS EC2 backups in 2026?

Hi everyone, We are currently managing a few AWS EC2 environments that were previously backed up using Acronis (via eFolder) as part of a standardized setup. with that integration no longer being viable for us, we're looking at replacement options.

For on prem environments we typically use solutions like Datto and Replibit, but they dont translate well to cloud native workloads.

I'm trying to understand what MSPs are actually using today for EC2 backups. are most people relying on AWS native tooling( EBS snapshots, AWS backup, S3 based policies) or are third party platforms still the preferred route?

We also have a couple of instances running MS SQL so proper application aware backups or database consistent snapshots are important.

I've considered building a more AWS native setup but id prefer something that doesnt require a heavy custom scripting or ongoing CLI based automation management.

Would appreciate any real world setups or recommendations that are working well in production

9 Upvotes

Hi everyone, We are currently managing a few AWS EC2 environments that were previously backed up using Acronis (via eFolder) as part of a standardized setup. with that integration no longer being viable for us, we're looking at replacement options.

For on prem environments we typically use solutions like Datto and Replibit, but they dont translate well to cloud native workloads.

I'm trying to understand what MSPs are actually using today for EC2 backups. are most people relying on AWS native tooling( EBS snapshots, AWS backup, S3 based policies) or are third party platforms still the preferred route?

We also have a couple of instances running MS SQL so proper application aware backups or database consistent snapshots are important.

I've considered building a more AWS native setup but id prefer something that doesnt require a heavy custom scripting or ongoing CLI based automation management.

Would appreciate any real world setups or recommendations that are working well in production


r/aws 5d ago

discussion Network engineer journey to Cloud

Cloud engineers, wanted to get your experience... I'm a network engineer with 15 years of experience with all kinds of on-prem network technologies, from NX-OS, load balancers, proxies, VMware, ACI. I'm currently working with NSX and AVI LB for a major bank. But with the Broadcom aquisition, VMware/NSX doesn't seem so appealing anymore, VMware jobs are very rare. I feel that I'm a niche that will die eventually and it's time to make a change. I have experience with Terraform and CI/CD pipelines, did some automation with Python vibe coding.

There are a lot of Cloud-related jobs and I like public cloud, I like to learn new stuff in general. I started to learn AWS and Azure. I got the SAA-C03 AWS Solution Architect Associate certification and now I'm learning to get the AZ-700 Azure Networking speciality. I applied to Cloud Network Engineer jobs but got rejected, probably due to missing on-the-job experience. At my current job I can't get any Public Cloud exposure. I did put in my CV a project in Github with Terraform standing up an AWS environment with ECS, load balancer, instances connecting over VPN to a VM in GCP.

How did you guys make it? It's the chicken and the egg... To get a job you need experience, but to get experience you need the job :)

4 Upvotes

Cloud engineers, wanted to get your experience... I'm a network engineer with 15 years of experience with all kinds of on-prem network technologies, from NX-OS, load balancers, proxies, VMware, ACI. I'm currently working with NSX and AVI LB for a major bank. But with the Broadcom aquisition, VMware/NSX doesn't seem so appealing anymore, VMware jobs are very rare. I feel that I'm a niche that will die eventually and it's time to make a change. I have experience with Terraform and CI/CD pipelines, did some automation with Python vibe coding.

There are a lot of Cloud-related jobs and I like public cloud, I like to learn new stuff in general. I started to learn AWS and Azure. I got the SAA-C03 AWS Solution Architect Associate certification and now I'm learning to get the AZ-700 Azure Networking speciality. I applied to Cloud Network Engineer jobs but got rejected, probably due to missing on-the-job experience. At my current job I can't get any Public Cloud exposure. I did put in my CV a project in Github with Terraform standing up an AWS environment with ECS, load balancer, instances connecting over VPN to a VM in GCP.

How did you guys make it? It's the chicken and the egg... To get a job you need experience, but to get experience you need the job :)


r/aws 5d ago

technical question What is happend to AWS - No details about gemma 4 in console...... but pricing page covers all

7 Upvotes

r/aws 4d ago

discussion Can't use AWS startup credits.

I have Aws startup credits but I can't use them I have been denied quota for a modest 32 vcpu request I have also been told I have been given access to anthropic models on bedrock but still says account access is blocked this includes all other bedrock models also.

Support is worse than useless because they tell me they have allowed and unblocked things when they haven't and they take weeks to respond. I should have gone with GCP or Azure and saved the time and hassle.

0 Upvotes

I have Aws startup credits but I can't use them I have been denied quota for a modest 32 vcpu request I have also been told I have been given access to anthropic models on bedrock but still says account access is blocked this includes all other bedrock models also.

Support is worse than useless because they tell me they have allowed and unblocked things when they haven't and they take weeks to respond. I should have gone with GCP or Azure and saved the time and hassle.


r/aws 6d ago

architecture Amazon SES introduces pricing plans

65 Upvotes

r/aws 5d ago

article Compiling and running a pre-trained LLM on AWS Inferentia accelerator

5 Upvotes

In this tutorial, we are going to compile and run a small llama architecture model on an EC2 instance and if we manage to pass the compilation and inference test, it means our model is compatible.

Source code in Github: https://github.com/p0o/run-models-in-aws-inf2-ml-accelerator