# AWS Service - S3

**S3**:

* **Buckets**
    
    * Amazon S3 allows people to store objects (files) in “buckets” (directories)
        
    * Buckets must have a globally unique name (across all regions all accounts)
        
    * Buckets are defined at the region level
        
    * Note: S3 looks like a global service but buckets are created in a region
        
    * Naming convention
        
        * No uppercase, No underscore
            
        * 3-63 characters long
            
        * Not an IP
            
        * Must start with lowercase letter or number
            
        * Must NOT start with the prefix `xn--`
            
        * Must NOT end with the suffix `-s3alias`
            
        * Doc link - [Naming convention](https://docs.aws.amazon.com/AmazonS3/latest/userguide/bucketnamingrules.html)
            
* **Objects**
    
    * Objects (files) have a Key
        
    * The key is the full path:
        
        * s3://my-bucket/my\_file.txt
            
        * s3://my-bucket/my\_folder1/another\_folder/my\_file.txt
            
    * The key is composed of prefix + object name
        
        * s3://my-bucket/my\_folder1/another\_folder/my\_file.txt
            
        * Here:
            
            * Bucket is `my-bucket`
                
            * Prefix is `my_folder1/another_folder`
                
            * Object is `my_file.txt`
                
    * There’s no concept of “directories” within buckets
        
    * Just keys with very long names that contain slashes (“/”)
        
    * Object values are the content of the body:
        
        * Max object Size is 5TB
            
        * If uploading more than 5GB, must use “multi-part upload”
            
    * Metadata (list of text key / value pairs - system or user metadata)
        
    * Tags (Unicode key / value pair - up to 10) - useful for security / lifecycle
        
    * Version ID (if versioning is enabled)
        
* **Security**
    
    * User-Based
        
        * Through IAM policies
            
            * which API calls should be allowed for a specific user from IAM
                
    * Resource-Based
        
        * Bucket Policies - bucket wide rules from the S3 console - allows cross account
            
        * Object Access Control List (ACL) - finer grain - can be disabled
            
        * Bucket Access Control List (ACL) - less common - can be disabled
            
    * Encryption
        
        * encrypt objects in Amazon S3 using encryption keys
            
* **S3 Bucket Policies**
    
    * JSON based policies - looks similar to IAM policies
        
    * Usecase:
        
        * Grant public access to the bucket
            
        * Force objects to be encrypted at upload
            
        * Grant access to another account (Cross Account)
            

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726600817888/bbe35f73-9436-4cd9-aed8-f59f6bf75c91.png align="center")

* **Bucket settings for Block Public Access**
    

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726600903536/29ca3db8-3471-49f8-adc5-d1acc1874ffe.png align="center")

* **Static Website Hosting**
    
    * S3 can host static websites and have them accessible on the Internet
        
    * The website URL will be either of the below depending on the region
        
        * http://bucket-name.s3-website-aws-region.amazonaws.com
            
        * http://bucket-name.s3-website.aws-region.amazonaws.com
            
        * the only diff is `website-` and `website.` in the above two pattern
            
* **Versioning**
    
    * You can version your files in Amazon S3
        
    * It is enabled at the bucket level
        
    * Same key overwrite will change the “version”: 1, 2, 3, etc
        
    * It is best practice to version your buckets
        
        * Protect against unintended deletes (ability to restore a version)
            
        * Roll back to previous version
            
    * Note:
        
        * Any file that is not versioned prior to enabling versioning will have version `null`
            
        * Suspending versioning does not delete the previous versions
            
* **Replication**
    
    * Must enable Versioning in source and destination buckets
        
    * CRR - Cross-Region Replication
        
    * SRR - Same-Region Replication
        
    * Buckets can be in different AWS accounts
        
    * Copying is asynchronous
        
    * Must give proper IAM permissions to S3
        
    * Usecase:
        
        * CRR - compliance, lower latency access, replication across account
            
        * SRR - log aggregation, live replication between envs
            
    * After you enable Replication, only new objects are replicated
        
    * Optionally, you can replicate existing objects using S3 Batch Replication
        
        * Replicates existing objects and objects that failed replication
            
    * For `Delete` operations
        
        * We could replicate delete markers from source to target (optional setting)
            
        * Deletions with a version ID are not replicated (to avoid malicious deletes)
            
    * There is no “chaining” of replication
        
        * If bucket 1 has replication into bucket 2, which has replication into bucket 3
            
        * Then objects created in bucket 1 are not replicated to bucket 3
            
* **Durability and Availability**
    
    * Durability:
        
        * High durability (99.999999999%, 11 9’s) of objects across multiple AZ
            
        * If you store 10,000,000 objects with Amazon S3, you can on average expect to incur a loss of a single object once every 10,000 years
            
        * Same for all storage classes
            
    * Availability:
        
        * Measures how readily available a service is
            
        * Varies depending on storage class
            
        * Example: S3 standard has 99.99% availability = not available 53 minutes a year
            
* **Storage Classes**
    
    * Amazon S3 Standard - General Purpose
        
    * Amazon S3 Standard-Infrequent Access (IA)
        
    * Amazon S3 One Zone-Infrequent Access
        
    * Amazon S3 Glacier Instant Retrieval
        
    * Amazon S3 Glacier Flexible Retrieval
        
    * Amazon S3 Glacier Deep Archive
        
    * Amazon S3 Intelligent Tiering
        
* **S3 Standard - General Purpose**
    
    * 99.99% Availability
        
    * Used for frequently accessed data
        
    * Low latency and high throughput
        
    * Sustain 2 concurrent facility failures
        
    * Use Cases: Big Data analytics, mobile & gaming applications, content distribution
        
* **S3 Storage Classes - Infrequent Access**
    
    * For data that is less frequently accessed, but requires rapid access when needed
        
    * Lower cost than S3 Standard
        
    * Amazon S3 Standard-Infrequent Access (S3 Standard-IA)
        
        * 99.9% Availability
            
        * Use cases: Disaster Recovery, backups
            
    * Amazon S3 One Zone-Infrequent Access (S3 One Zone-IA)
        
        * High durability (99.999999999%) in a single AZ; data lost when AZ is destroyed
            
        * 99.5% Availability
            
        * Use Cases: Storing secondary backup copies of on-premises data, or data you can recreate
            
* **S3 Glacier Storage Classes**
    
    * Low-cost object storage meant for archiving / backup
        
    * Pricing: price for storage + object retrieval cost
        
    * Amazon S3 Glacier Instant Retrieval
        
        * Millisecond retrieval, great for data accessed once a quarter
            
        * Minimum storage duration of 90 days
            
    * Amazon S3 Glacier Flexible Retrieval
        
        * Expedited (1 to 5 minutes), Standard (3 to 5 hours), Bulk (5 to 12 hours) - free
            
        * Minimum storage duration of 90 days
            
    * Amazon S3 Glacier Deep Archive - for long term storage
        
        * Standard (12 hours), Bulk (48 hours)
            
        * Minimum storage duration of 180 days
            
* **S3 Intelligent-Tiering**
    
    * Moves objects automatically between Access Tiers based on usage
        
    * Small monthly monitoring and auto-tiering fee
        
    * There are no retrieval charges in S3 Intelligent-Tiering
        
    * Frequent Access tier (automatic): default tier
        
    * Infrequent Access tier (automatic): objects not accessed for 30 days
        
    * Archive Instant Access tier (automatic): objects not accessed for 90 days
        
    * Archive Access tier (optional): configurable from 90 days to 700+ days
        
    * Deep Archive Access tier (optional): config. from 180 days to 700+ days
        
* **S3 Storage Classes Comparison**
    

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726639449366/78e25b8a-4a4f-462b-998d-a439ea7851b6.png align="center")

* **S3 Storage Classes - Price Comparison**
    

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726639501945/d248c5b3-c1a4-4895-b388-f3af25c63de6.png align="center")

* **Moving between Storage Classes**
    
    * You can transition objects between storage classes
        
    * For infrequently accessed object, move them to Standard IA
        
    * For archive objects that you don’t need fast access to, move them to Glacier or Glacier Deep Archive
        
    * Moving objects can be automated using a Lifecycle Rules
        
* **Lifecycle Rules**
    
    * Transition Actions
        
        * Configure objects to transition to another storage class
            
        * Move objects to Standard IA class 60 days after creation
            
        * Move to Glacier for archiving after 6 months
            
    * Expiration actions
        
        * Configure objects to delete after some time
            
        * Access log files can be set to delete after a 365 days
            
        * Can be used to delete old versions of files (if versioning is enabled)
            
        * Can be used to delete incomplete Multi-Part uploads
            
    * Rules can be created for a certain prefix, certain objects Tags
        
* **Storage Class Analysis**
    
    * Help you decide when to transition objects to the right storage class
        
    * Recommendations for Standard and Standard IA
        
        * Does not work for One-Zone IA or Glacier
            
    * Report is updated daily
        
    * 24 to 48 hours to start seeing data analysis
        
* **Requester Pays**
    
    * In general, bucket owners pay for all Amazon S3 storage and data transfer costs associated with their bucket
        
    * With Requester Pays buckets, the requester instead of the bucket owner pays the cost of the request and the data download from the bucket
        
    * Helpful when you want to share large datasets with other accounts
        
    * The requester must be authenticated in AWS (cannot be anonymous)
        
* **Event Notifications**
    
    * Based on S3 events downstream services could be triggered
        
    * S3:ObjectCreated, S3:ObjectRemoved, S3:ObjectRestore, S3:Replication
        
    * Can create as many “S3 events” as desired
        
    * S3 event notifications typically deliver events in seconds but can sometimes take a minute or longer
        

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726687790611/8e18ca9c-5acf-4ab9-9291-37b8aa972f18.png align="center")

* **S3 Event Notifications with Amazon EventBridge**
    
    * Advanced filtering options with JSON rules
        
    * Multiple Destinations - Lambda, SNS, Step function, etc
        
    * EventBridge Capabilities - Archive, Replay Events, Reliable delivery
        

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726687838019/0999ef33-b497-4d07-b025-cc7a6943c2a4.png align="center")

* **Baseline Performance**
    
    * Amazon S3 automatically scales to high request rates, latency 100-200 ms
        
    * We could achieve at least 3,500 PUT/COPY/POST/DELETE or 5,500 GET/HEAD requests per second per prefix in a bucket
        
    * There are no limits to the number of prefixes in a bucket
        
    * Tip: Spreading reads across prefix evenly could give higher reads per second
        
        * Ex: Prefix1 ; Prefix 2
            
        * So we basically get 2 × 5500 = 11,000 api per second indirectly by abusing the fact that there are no limits to no of prefix and per prefix we could go as high as 5,500 per sec
            
* **S3 Performance**
    
    * Multi-Part upload
        
        * Recommended for files &gt; 100MB, must use for files &gt; 5GB
            
        * Can help parallelize uploads (speed up transfers)
            
    * S3 Transfer Acceleration
        
        * Increase transfer speed by transferring file to an AWS edge location which will forward the data to the S3 bucket in the target region
            
        * Compatible with multi-part upload
            
* **S3 Byte-Range Fetches**
    
    * Parallelize GETs by requesting specific byte ranges
        
    * Better resilience in case of failures
        
    * Can be used to speed up downloads
        
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726688348367/a15173bd-c6b1-4268-a929-c3f2fd9c0f37.png align="center")
    
    * Can be used to retrieve only partial data (for example the head of a file)
        

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726688367347/6ac453d7-864e-4a22-9505-832cdee82bb7.png align="center")

* **S3 Select & Glacier Select**
    
    * Retrieve less data using SQL by performing server-side filtering
        
    * Can filter by rows & columns (simple SQL statements)
        
    * Less network transfer, less CPU cost client-side
        

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726688445126/f10f34a8-e980-4131-9172-5d3a208e21f3.png align="center")

* **S3 Batch Operations**
    
    * Perform bulk operations on existing S3 objects with a single request, example:
        
        * Modify object metadata and properties
            
        * Copy objects between S3 buckets
            
        * Encrypt un-encrypted objects
            
        * Modify ACLs, tags
            
        * Restore objects from S3 Glacier
            
        * Invoke Lambda function to perform custom action on each object
            
    * A job consists of a list of objects, the action to perform, and optional parameters
        
    * S3 Batch Operations manages retries, tracks progress, sends completion notifications, generate reports
        
    * You can use S3 Inventory to get object list and use S3 Select to filter your objects
        

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726775161120/36f2dccf-4267-4b0a-8b85-8402af0ace8b.png align="center")

* **S3 Object Encryption**
    
    * Server-Side Encryption with Amazon S3-Managed Keys (SSE-S3)
        
    * Server-Side Encryption with KMS Keys stored in AWS KMS (SSE-KMS)
        
    * Server-Side Encryption with Customer-Provided Keys (SSE-C)
        
    * Client-Side Encryption
        
* **S3 Encryption - SSE -S3**
    
    * Encryption using keys handled, managed, and owned by AWS
        
    * Object is encrypted server-side
        
    * Encryption type is AES-256
        
    * Must set header "x-amz-server-side-encryption": "AES256”
        
    * Enabled by default for new buckets & new objects
        
* **S3 Encryption - SSE-KMS**
    
    * Encryption using keys handled and managed by AWS KMS (Key Management Service)
        
    * Better user control and audit key usage tracking using CloudTrail
        
    * Object is encrypted server side
        
    * Must set header "x-amz-server-side-encryption": "aws:kms"
        
    * Beware there is quota limit to KMS API rates
        
* **Amazon S3 Encryption - SSE-C**
    
    * Server-Side Encryption using keys fully managed by the customer outside of AWS
        
    * Amazon S3 does not store the encryption key you provide
        
    * HTTPS must be used
        
    * Encryption key must provided in HTTP headers, for every HTTP request made
        
* **Amazon S3 Encryption - Client-Side Encryption**
    
    * Use client libraries such as Amazon S3 Client-Side Encryption Library
        
    * Clients must encrypt data themselves before sending to Amazon S3
        
    * Clients must decrypt data themselves when retrieving from Amazon S3
        
    * Customer fully manages the keys and encryption cycle
        
* **Encryption in transit (SSL/TLS)**
    
    * Encryption in flight is also called SSL/TLS
        
    * Amazon S3 exposes two endpoints:
        
        * HTTP Endpoint - non encrypted
            
        * HTTPS Endpoint - encryption in flight
            
    * HTTPS is mandatory for SSE-C
        
    * Most clients would use the HTTPS endpoint by default
        
    * We could force HTTPS by adding those condition in bucket policy
        
* **CORS**
    
    * CORS - Cross-Origin Resource Sharing
        
    * Origin = scheme (protocol) + host (domain) + port
        
        * ex: https://www.example.com
            
        * Protocol - HTTPS
            
        * Domain - www.example.com
            
        * Port - HTTPS (443) ; HTTP (80)
            
    * Web Browser based mechanism to allow requests to other origins while visiting the main origin
        
    * Same origin: `http://example.com/page1` and `http://example.com/page2`
        
    * Different origins: `http://www.example.com` and `http://other.example.com`
        
    * The requests won’t be fulfilled unless the other origin allows for the requests, using CORS Headers (example: Access-Control-Allow-Origin)
        
    * If a client makes a cross-origin request on our S3 bucket, we need to enable the correct CORS headers
        
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726776206776/4418dc55-0ebc-474e-8ad5-dc651b08e938.png align="center")
    
* **S3 - MFA Delete**
    
    * MFA (Multi-Factor Authentication) – force users to generate a code on a device (usually a mobile phone or hardware) before doing important operations on S3
        
    * MFA will be required to:
        
        * Permanently delete an object version
            
        * Suspend Versioning on the bucket
            
    * MFA won’t be required to:
        
        * Enable Versioning
            
        * List deleted versions
            
    * To use MFA Delete, Versioning must be enabled on the bucke
        
    * Only the bucket owner (root account) can enable/disable MFA Delete
        
* **S3 Access Logs**
    
    * For audit purpose, you may want to log all access to S3 buckets
        
    * Any request made to S3, from any account, authorized or denied, will be logged into another S3 bucket
        
    * That data can be analyzed using data analysis tools
        
    * The target logging bucket must be in the same AWS region
        
    * Do not set your logging bucket to be the monitored bucket
        
    * It will create a logging loop, and your bucket will grow exponentially - you will pay the price for your negligence or some bad screwing you up
        
* **S3 - Pre-Signed URLs**
    
    * Generate pre-signed URLs using the S3 Console, AWS CLI or SDK
        
    * URL Expiration
        
        * S3 Console - 1 min up to 720 mins (12 hours)
            
        * AWS CLI - configure expiration with --expires-in parameter in seconds (default 3600 secs, max. 604800 secs ~ 168 hours)
            
    * Users given a pre-signed URL inherit the permissions of the user that generated the URL for GET / PUT
        
    * Examples:
        
        * Allow only logged-in users to download a premium video from your S3 bucket
            
        * Allow an ever-changing list of users to download files by generating URLs dynamically
            
        * Allow temporarily a user to upload a file to a precise location in your S3 bucket
            

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726776548032/fa16cb09-8c9a-4992-a1d8-1e5f2cf3ffb8.png align="center")

* **S3 Glacier Vault Lock**
    
    * Adopt a WORM (Write Once Read Many) model
        
    * Create a Vault Lock Policy
        
    * Lock the policy for future edits - can no longer be changed or deleted
        
    * Helpful for compliance and data retention
        
* **S3 Object Lock**
    
    * Versioning must be enabled
        
    * Adopt a WORM (Write Once Read Many) model
        
    * Block an object version deletion for a specified amount of time
        
    * Retention mode - Compliance:
        
        * Object versions can't be overwritten or deleted by any user, including the root user
            
        * Objects retention modes can't be changed, and retention periods can't be shortened
            
    * Retention mode - Governance:
        
        * Most users can't overwrite or delete an object version or alter its lock settings
            
        * Some users have special permissions to change the retention or delete the object
            
    * Retention Period: protect the object for a fixed period, it can be extended
        
    * Legal Hold:
        
        * Protect the object indefinitely, independent from retention period
            
        * Can be freely placed and removed using the s3:PutObjectLegalHold IAM permission
            
* **S3 - Access Points**
    
    * Access Points simplify security management for S3 Buckets
        
    * Each Access Point has
        
        * Own DNS name (Internet Origin or VPC Origin)
            
        * An access point policy similar to bucket policy to manage security at scale
            
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726776791655/bc69c560-b935-48e8-a73b-e689c4b09e4c.png align="center")
    
    * We can define the access point to be accessible only from within the VPC
        
    * You must create a VPC Endpoint to access the Access Point (Gateway or Interface Endpoint)
        
    * The VPC Endpoint Policy must allow access to the target bucket and Access Point
        

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726776831756/014d65c8-d939-4e43-b98a-e7c3ae7f0c26.png align="center")

* **S3 Object Lambda**
    
    * Use AWS Lambda Functions to change the object before it is retrieved by the caller application
        
    * Only one S3 bucket is needed, on top of which we create S3 Access Point and S3 Object Lambda Acces
        
    * Use Case:
        
        * Redacting personally identifiable information for analytics or non- production environment
            
        * Converting across data formats, such as converting XML to JSON
            
        * Resizing and watermarking images on the fly using caller-specific details, such as the user who requested the object
            

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1726776904175/49093c0d-a374-4f96-8d2e-4a14c5ed122b.png align="center")

**Relevant Doc:**

[https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html](https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html)

**Disclaimer:** This is a personal blog that might come in handy when I suffer from Dementia in future
