Skip to main content

Command Palette

Search for a command to run...

AWS Service - S3

Published
•13 min read•View as Markdown
B

I currently work

S3:

  • Buckets

    • Amazon S3 allows people to store objects (files) in “buckets” (directories)

    • Buckets must have a globally unique name (across all regions all accounts)

    • Buckets are defined at the region level

    • Note: S3 looks like a global service but buckets are created in a region

    • Naming convention

      • No uppercase, No underscore

      • 3-63 characters long

      • Not an IP

      • Must start with lowercase letter or number

      • Must NOT start with the prefix xn--

      • Must NOT end with the suffix -s3alias

      • Doc link - Naming convention

  • Objects

    • Objects (files) have a Key

    • The key is the full path:

      • s3://my-bucket/my_file.txt

      • s3://my-bucket/my_folder1/another_folder/my_file.txt

    • The key is composed of prefix + object name

      • s3://my-bucket/my_folder1/another_folder/my_file.txt

      • Here:

        • Bucket is my-bucket

        • Prefix is my_folder1/another_folder

        • Object is my_file.txt

    • There’s no concept of “directories” within buckets

    • Just keys with very long names that contain slashes (“/”)

    • Object values are the content of the body:

      • Max object Size is 5TB

      • If uploading more than 5GB, must use “multi-part upload”

    • Metadata (list of text key / value pairs - system or user metadata)

    • Tags (Unicode key / value pair - up to 10) - useful for security / lifecycle

    • Version ID (if versioning is enabled)

  • Security

    • User-Based

      • Through IAM policies

        • which API calls should be allowed for a specific user from IAM
    • Resource-Based

      • Bucket Policies - bucket wide rules from the S3 console - allows cross account

      • Object Access Control List (ACL) - finer grain - can be disabled

      • Bucket Access Control List (ACL) - less common - can be disabled

    • Encryption

      • encrypt objects in Amazon S3 using encryption keys
  • S3 Bucket Policies

    • JSON based policies - looks similar to IAM policies

    • Usecase:

      • Grant public access to the bucket

      • Force objects to be encrypted at upload

      • Grant access to another account (Cross Account)

  • Bucket settings for Block Public Access

  • Static Website Hosting

    • S3 can host static websites and have them accessible on the Internet

    • The website URL will be either of the below depending on the region

      • http://bucket-name.s3-website-aws-region.amazonaws.com

      • http://bucket-name.s3-website.aws-region.amazonaws.com

      • the only diff is website- and website. in the above two pattern

  • Versioning

    • You can version your files in Amazon S3

    • It is enabled at the bucket level

    • Same key overwrite will change the “version”: 1, 2, 3, etc

    • It is best practice to version your buckets

      • Protect against unintended deletes (ability to restore a version)

      • Roll back to previous version

    • Note:

      • Any file that is not versioned prior to enabling versioning will have version null

      • Suspending versioning does not delete the previous versions

  • Replication

    • Must enable Versioning in source and destination buckets

    • CRR - Cross-Region Replication

    • SRR - Same-Region Replication

    • Buckets can be in different AWS accounts

    • Copying is asynchronous

    • Must give proper IAM permissions to S3

    • Usecase:

      • CRR - compliance, lower latency access, replication across account

      • SRR - log aggregation, live replication between envs

    • After you enable Replication, only new objects are replicated

    • Optionally, you can replicate existing objects using S3 Batch Replication

      • Replicates existing objects and objects that failed replication
    • For Delete operations

      • We could replicate delete markers from source to target (optional setting)

      • Deletions with a version ID are not replicated (to avoid malicious deletes)

    • There is no “chaining” of replication

      • If bucket 1 has replication into bucket 2, which has replication into bucket 3

      • Then objects created in bucket 1 are not replicated to bucket 3

  • Durability and Availability

    • Durability:

      • High durability (99.999999999%, 11 9’s) of objects across multiple AZ

      • If you store 10,000,000 objects with Amazon S3, you can on average expect to incur a loss of a single object once every 10,000 years

      • Same for all storage classes

    • Availability:

      • Measures how readily available a service is

      • Varies depending on storage class

      • Example: S3 standard has 99.99% availability = not available 53 minutes a year

  • Storage Classes

    • Amazon S3 Standard - General Purpose

    • Amazon S3 Standard-Infrequent Access (IA)

    • Amazon S3 One Zone-Infrequent Access

    • Amazon S3 Glacier Instant Retrieval

    • Amazon S3 Glacier Flexible Retrieval

    • Amazon S3 Glacier Deep Archive

    • Amazon S3 Intelligent Tiering

  • S3 Standard - General Purpose

    • 99.99% Availability

    • Used for frequently accessed data

    • Low latency and high throughput

    • Sustain 2 concurrent facility failures

    • Use Cases: Big Data analytics, mobile & gaming applications, content distribution

  • S3 Storage Classes - Infrequent Access

    • For data that is less frequently accessed, but requires rapid access when needed

    • Lower cost than S3 Standard

    • Amazon S3 Standard-Infrequent Access (S3 Standard-IA)

      • 99.9% Availability

      • Use cases: Disaster Recovery, backups

    • Amazon S3 One Zone-Infrequent Access (S3 One Zone-IA)

      • High durability (99.999999999%) in a single AZ; data lost when AZ is destroyed

      • 99.5% Availability

      • Use Cases: Storing secondary backup copies of on-premises data, or data you can recreate

  • S3 Glacier Storage Classes

    • Low-cost object storage meant for archiving / backup

    • Pricing: price for storage + object retrieval cost

    • Amazon S3 Glacier Instant Retrieval

      • Millisecond retrieval, great for data accessed once a quarter

      • Minimum storage duration of 90 days

    • Amazon S3 Glacier Flexible Retrieval

      • Expedited (1 to 5 minutes), Standard (3 to 5 hours), Bulk (5 to 12 hours) - free

      • Minimum storage duration of 90 days

    • Amazon S3 Glacier Deep Archive - for long term storage

      • Standard (12 hours), Bulk (48 hours)

      • Minimum storage duration of 180 days

  • S3 Intelligent-Tiering

    • Moves objects automatically between Access Tiers based on usage

    • Small monthly monitoring and auto-tiering fee

    • There are no retrieval charges in S3 Intelligent-Tiering

    • Frequent Access tier (automatic): default tier

    • Infrequent Access tier (automatic): objects not accessed for 30 days

    • Archive Instant Access tier (automatic): objects not accessed for 90 days

    • Archive Access tier (optional): configurable from 90 days to 700+ days

    • Deep Archive Access tier (optional): config. from 180 days to 700+ days

  • S3 Storage Classes Comparison

  • S3 Storage Classes - Price Comparison

  • Moving between Storage Classes

    • You can transition objects between storage classes

    • For infrequently accessed object, move them to Standard IA

    • For archive objects that you don’t need fast access to, move them to Glacier or Glacier Deep Archive

    • Moving objects can be automated using a Lifecycle Rules

  • Lifecycle Rules

    • Transition Actions

      • Configure objects to transition to another storage class

      • Move objects to Standard IA class 60 days after creation

      • Move to Glacier for archiving after 6 months

    • Expiration actions

      • Configure objects to delete after some time

      • Access log files can be set to delete after a 365 days

      • Can be used to delete old versions of files (if versioning is enabled)

      • Can be used to delete incomplete Multi-Part uploads

    • Rules can be created for a certain prefix, certain objects Tags

  • Storage Class Analysis

    • Help you decide when to transition objects to the right storage class

    • Recommendations for Standard and Standard IA

      • Does not work for One-Zone IA or Glacier
    • Report is updated daily

    • 24 to 48 hours to start seeing data analysis

  • Requester Pays

    • In general, bucket owners pay for all Amazon S3 storage and data transfer costs associated with their bucket

    • With Requester Pays buckets, the requester instead of the bucket owner pays the cost of the request and the data download from the bucket

    • Helpful when you want to share large datasets with other accounts

    • The requester must be authenticated in AWS (cannot be anonymous)

  • Event Notifications

    • Based on S3 events downstream services could be triggered

    • S3:ObjectCreated, S3:ObjectRemoved, S3:ObjectRestore, S3:Replication

    • Can create as many “S3 events” as desired

    • S3 event notifications typically deliver events in seconds but can sometimes take a minute or longer

  • S3 Event Notifications with Amazon EventBridge

    • Advanced filtering options with JSON rules

    • Multiple Destinations - Lambda, SNS, Step function, etc

    • EventBridge Capabilities - Archive, Replay Events, Reliable delivery

  • Baseline Performance

    • Amazon S3 automatically scales to high request rates, latency 100-200 ms

    • We could achieve at least 3,500 PUT/COPY/POST/DELETE or 5,500 GET/HEAD requests per second per prefix in a bucket

    • There are no limits to the number of prefixes in a bucket

    • Tip: Spreading reads across prefix evenly could give higher reads per second

      • Ex: Prefix1 ; Prefix 2

      • So we basically get 2 × 5500 = 11,000 api per second indirectly by abusing the fact that there are no limits to no of prefix and per prefix we could go as high as 5,500 per sec

  • S3 Performance

    • Multi-Part upload

      • Recommended for files > 100MB, must use for files > 5GB

      • Can help parallelize uploads (speed up transfers)

    • S3 Transfer Acceleration

      • Increase transfer speed by transferring file to an AWS edge location which will forward the data to the S3 bucket in the target region

      • Compatible with multi-part upload

  • S3 Byte-Range Fetches

    • Parallelize GETs by requesting specific byte ranges

    • Better resilience in case of failures

    • Can be used to speed up downloads

  • Can be used to retrieve only partial data (for example the head of a file)

  • S3 Select & Glacier Select

    • Retrieve less data using SQL by performing server-side filtering

    • Can filter by rows & columns (simple SQL statements)

    • Less network transfer, less CPU cost client-side

  • S3 Batch Operations

    • Perform bulk operations on existing S3 objects with a single request, example:

      • Modify object metadata and properties

      • Copy objects between S3 buckets

      • Encrypt un-encrypted objects

      • Modify ACLs, tags

      • Restore objects from S3 Glacier

      • Invoke Lambda function to perform custom action on each object

    • A job consists of a list of objects, the action to perform, and optional parameters

    • S3 Batch Operations manages retries, tracks progress, sends completion notifications, generate reports

    • You can use S3 Inventory to get object list and use S3 Select to filter your objects

  • S3 Object Encryption

    • Server-Side Encryption with Amazon S3-Managed Keys (SSE-S3)

    • Server-Side Encryption with KMS Keys stored in AWS KMS (SSE-KMS)

    • Server-Side Encryption with Customer-Provided Keys (SSE-C)

    • Client-Side Encryption

  • S3 Encryption - SSE -S3

    • Encryption using keys handled, managed, and owned by AWS

    • Object is encrypted server-side

    • Encryption type is AES-256

    • Must set header "x-amz-server-side-encryption": "AES256”

    • Enabled by default for new buckets & new objects

  • S3 Encryption - SSE-KMS

    • Encryption using keys handled and managed by AWS KMS (Key Management Service)

    • Better user control and audit key usage tracking using CloudTrail

    • Object is encrypted server side

    • Must set header "x-amz-server-side-encryption": "aws:kms"

    • Beware there is quota limit to KMS API rates

  • Amazon S3 Encryption - SSE-C

    • Server-Side Encryption using keys fully managed by the customer outside of AWS

    • Amazon S3 does not store the encryption key you provide

    • HTTPS must be used

    • Encryption key must provided in HTTP headers, for every HTTP request made

  • Amazon S3 Encryption - Client-Side Encryption

    • Use client libraries such as Amazon S3 Client-Side Encryption Library

    • Clients must encrypt data themselves before sending to Amazon S3

    • Clients must decrypt data themselves when retrieving from Amazon S3

    • Customer fully manages the keys and encryption cycle

  • Encryption in transit (SSL/TLS)

    • Encryption in flight is also called SSL/TLS

    • Amazon S3 exposes two endpoints:

      • HTTP Endpoint - non encrypted

      • HTTPS Endpoint - encryption in flight

    • HTTPS is mandatory for SSE-C

    • Most clients would use the HTTPS endpoint by default

    • We could force HTTPS by adding those condition in bucket policy

  • CORS

    • CORS - Cross-Origin Resource Sharing

    • Origin = scheme (protocol) + host (domain) + port

      • ex: https://www.example.com

      • Protocol - HTTPS

      • Domain - www.example.com

      • Port - HTTPS (443) ; HTTP (80)

    • Web Browser based mechanism to allow requests to other origins while visiting the main origin

    • Same origin: http://example.com/page1 and http://example.com/page2

    • Different origins: http://www.example.com and http://other.example.com

    • The requests won’t be fulfilled unless the other origin allows for the requests, using CORS Headers (example: Access-Control-Allow-Origin)

    • If a client makes a cross-origin request on our S3 bucket, we need to enable the correct CORS headers

  • S3 - MFA Delete

    • MFA (Multi-Factor Authentication) – force users to generate a code on a device (usually a mobile phone or hardware) before doing important operations on S3

    • MFA will be required to:

      • Permanently delete an object version

      • Suspend Versioning on the bucket

    • MFA won’t be required to:

      • Enable Versioning

      • List deleted versions

    • To use MFA Delete, Versioning must be enabled on the bucke

    • Only the bucket owner (root account) can enable/disable MFA Delete

  • S3 Access Logs

    • For audit purpose, you may want to log all access to S3 buckets

    • Any request made to S3, from any account, authorized or denied, will be logged into another S3 bucket

    • That data can be analyzed using data analysis tools

    • The target logging bucket must be in the same AWS region

    • Do not set your logging bucket to be the monitored bucket

    • It will create a logging loop, and your bucket will grow exponentially - you will pay the price for your negligence or some bad screwing you up

  • S3 - Pre-Signed URLs

    • Generate pre-signed URLs using the S3 Console, AWS CLI or SDK

    • URL Expiration

      • S3 Console - 1 min up to 720 mins (12 hours)

      • AWS CLI - configure expiration with --expires-in parameter in seconds (default 3600 secs, max. 604800 secs ~ 168 hours)

    • Users given a pre-signed URL inherit the permissions of the user that generated the URL for GET / PUT

    • Examples:

      • Allow only logged-in users to download a premium video from your S3 bucket

      • Allow an ever-changing list of users to download files by generating URLs dynamically

      • Allow temporarily a user to upload a file to a precise location in your S3 bucket

  • S3 Glacier Vault Lock

    • Adopt a WORM (Write Once Read Many) model

    • Create a Vault Lock Policy

    • Lock the policy for future edits - can no longer be changed or deleted

    • Helpful for compliance and data retention

  • S3 Object Lock

    • Versioning must be enabled

    • Adopt a WORM (Write Once Read Many) model

    • Block an object version deletion for a specified amount of time

    • Retention mode - Compliance:

      • Object versions can't be overwritten or deleted by any user, including the root user

      • Objects retention modes can't be changed, and retention periods can't be shortened

    • Retention mode - Governance:

      • Most users can't overwrite or delete an object version or alter its lock settings

      • Some users have special permissions to change the retention or delete the object

    • Retention Period: protect the object for a fixed period, it can be extended

    • Legal Hold:

      • Protect the object indefinitely, independent from retention period

      • Can be freely placed and removed using the s3:PutObjectLegalHold IAM permission

  • S3 - Access Points

    • Access Points simplify security management for S3 Buckets

    • Each Access Point has

      • Own DNS name (Internet Origin or VPC Origin)

      • An access point policy similar to bucket policy to manage security at scale

  • We can define the access point to be accessible only from within the VPC

  • You must create a VPC Endpoint to access the Access Point (Gateway or Interface Endpoint)

  • The VPC Endpoint Policy must allow access to the target bucket and Access Point

  • S3 Object Lambda

    • Use AWS Lambda Functions to change the object before it is retrieved by the caller application

    • Only one S3 bucket is needed, on top of which we create S3 Access Point and S3 Object Lambda Acces

    • Use Case:

      • Redacting personally identifiable information for analytics or non- production environment

      • Converting across data formats, such as converting XML to JSON

      • Resizing and watermarking images on the fly using caller-specific details, such as the user who requested the object

Relevant Doc:

https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html

Disclaimer: This is a personal blog that might come in handy when I suffer from Dementia in future

More from this blog