AI/ML, AWS, Cloud Computing

< 1 min

Advanced Prompt Optimization in Amazon Bedrock for Better AI Performance

Voiced by Amazon Polly

Overview

Amazon Bedrock Advanced Prompt Optimization is a new managed capability that helps developers improve prompts and migrate generative AI applications between foundation models. It can evaluate the original and optimized versions of a prompt across up to five models in the same optimization job.

Instead of rewriting a prompt based only on general prompt-engineering guidelines, the service uses example inputs, expected outputs, and evaluation criteria to guide the optimization process.

The result is a more systematic approach to prompt engineering. Teams can compare prompt quality, model performance, cost, and latency before deciding whether to introduce an optimized prompt or a different foundation model into production.

Empowering organizations to become ‘data driven’ enterprises with our Cloud experts.

  • Reduced infrastructure costs
  • Timely data-driven decisions
Get Started

Introduction

Prompts often begin as simple instructions. During application development, they gradually expand to include business rules, output formats, examples, safety instructions, tool descriptions, and exception handling.

After several months, a production prompt may contain hundreds or thousands of words. Changing even one section can improve one use case while unintentionally reducing quality somewhere else.

The challenge becomes more serious when an organization migrates to a new foundation model. A prompt written for one model may not behave in the same way with another model. Differences in instruction following, reasoning, formatting, tool use, and context interpretation can lead to unexpected regressions.

Amazon Bedrock Advanced Prompt Optimization introduces an automated, evaluation-driven workflow for this process. It tests the prompt against representative data, measures the responses using a selected evaluation method, rewrites the prompt, and repeats the cycle to improve performance.

Advanced Prompt Optimization

Advanced Prompt Optimization is a managed Amazon Bedrock tool for improving prompts using evaluation feedback.

A developer provides:

  • The original prompt template
  • Example input-variable values
  • Optional reference or ground-truth responses
  • One or more models to test
  • An evaluation metric or optimization instruction

Amazon Bedrock runs the prompt against the selected models, evaluates the responses, rewrites the prompt, and continues the optimization loop based on the evaluation results. The final output includes the original and optimized prompts, evaluation scores, estimated cost, and latency information.

How It Differs from Simple Prompt Optimization

Amazon Bedrock also offers simple prompt optimization. The simple option performs a quick heuristic rewrite on a single model and is best suited to relatively short prompts. It does not use an evaluation dataset and does not compare performance across multiple models.

Advanced Prompt Optimization is designed for more structured and production-oriented work. It uses examples and evaluation criteria, supports model comparison, and assesses whether the new prompt actually improves performance on the intended task.

A simple optimizer may rewrite “Summarize this document” into a clearer instruction. An advanced optimizer can test whether the rewritten prompt preserves important facts, follows the required format, avoids unsupported claims, and performs consistently across a representative evaluation set.

Compare Up to Five Models

One of the most useful features is the ability to optimize and compare a prompt across up to five inference models simultaneously.

During a model migration, a team can use its current model as the baseline and select up to four candidate models. Bedrock then evaluates the original and optimized prompts across the selected models.

This helps answer several practical questions:

  • Will the new model maintain the current level of accuracy?
  • Does it follow the expected output format?
  • Which model provides the best balance of quality and cost?
  • Does the optimized prompt reduce latency?
  • Are there specific test cases where the new model performs worse?

Without this comparison, organizations may select a model based on a few manually tested questions. Advanced Prompt Optimization encourages a broader, evidence-based decision.

Evaluation-Driven Feedback Loop

The core of the feature is its feedback loop.

Bedrock sends the prompt and evaluation samples to the selected model. It scores the model’s responses using the configured evaluation method. The optimizer then rewrites the prompt based on the score and repeats the process, steering the prompt toward better results.

This differs from asking an LLM to “make this prompt better.” Improvement is tied to a measurable definition of success.

For a text-to-SQL application, success may mean execution accuracy and valid SQL syntax. For a document-extraction workflow, it may mean an exact match against expected JSON. For a customer support assistant, it may involve accuracy, empathy, policy compliance, and tone.

Three Evaluation Options

Advanced Prompt Optimization supports three main ways to define prompt quality.

AWS Lambda Evaluation

An AWS Lambda function can implement deterministic scoring logic. This works well when the output can be programmatically tested.

For example, the function could measure exact-match accuracy, F1 score, SQL execution accuracy, required-field coverage, JSON validity, or compliance with a fixed output schema. It compares the model response with the reference answer and returns a numerical score.

This approach is particularly useful for extraction, classification, code generation, text-to-SQL, and structured-output applications.

LLM as a Judge

Some tasks cannot be evaluated through an exact comparison. A high-quality summary may use different wording from the reference while still preserving all important information.

For such cases, developers can define a custom evaluation rubric and use an Amazon Bedrock model as the judge. The rubric can describe the metrics, scoring scale, and conditions that a strong response must satisfy. The judge evaluates each response and returns a score with their assessment.

This method is suitable for summarization, reasoning explanations, marketing content, customer communication, and other open-ended generation tasks.

Natural-Language Steering Criteria

Teams that do not need a detailed scoring implementation can provide natural-language criteria. Examples might include:

“Use a professional but conversational tone.”

“Return valid JSON without additional commentary.”

“Do not answer beyond the supplied context.”

“Keep the response below 200 words.”

The default judge incorporates these criteria into its evaluation and uses them to guide the prompt-rewriting process.

Multimodal Prompt Optimization

Advanced Prompt Optimization is not limited to text-only tasks. It supports PDF, JPG, and PNG inputs within prompt templates.

This allows teams to optimize prompts for document analysis, image understanding, invoice extraction, form processing, visual inspection, and other multimodal workloads. Files can be referenced from Amazon S3 as part of the evaluation samples.

For example, an insurance application could provide sample claim documents and expected extraction results. Bedrock could then optimize the prompt for accurately identifying policy numbers, claim amounts, dates, and required supporting evidence.

Preparing an Optimization Job

Prompt templates and evaluation data are prepared in JSON Lines format. A job can contain multiple prompt templates, and each template can include its own input variables, reference responses, multimodal inputs, and evaluation configuration.

The dataset can be uploaded directly or imported from Amazon S3. An Amazon S3 output location is also selected for the optimized prompts and evaluation results.

A typical workflow is:

  1. Collect representative production questions.
  2. Add expected responses where available.
  3. Select a baseline and candidate models.
  4. Choose an evaluation approach.
  5. Run the optimization job.
  6. Review scores, cost, and latency.
  7. Test the optimized prompt against a separate validation dataset.
  8. Introduce the change through the normal application release process.

The optimization job can be created through the Amazon Bedrock console or programmatically using the CreateAdvancedPromptOptimizationJob API.

Conclusion

Amazon Bedrock Advanced Prompt Optimization turns prompt improvement from a largely manual activity into a measurable engineering process.

It allows teams to test prompts on real examples, compare multiple models, define task-specific evaluation criteria, identify regressions, and review quality alongside cost and latency. Its support for deterministic Lambda metrics, LLM-based evaluation, natural-language steering, and multimodal inputs makes it suitable for a wide range of enterprise generative AI applications.

Drop a query if you have any questions regarding Amazon Bedrock, and we will get back to you quickly.

Pioneers in Cloud Consulting & Migration Services

  • Reduced infrastructural costs
  • Accelerated application deployment
Get Started

About CloudThat

CloudThat is an award-winning company and the first in India to offer cloud training and consulting services worldwide. As an AWS Premier Tier Services Partner, AWS Advanced Training Partner, Microsoft Solutions Partner, and Google Cloud Platform Partner, CloudThat has empowered over 1.1 million professionals through 1000+ cloud certifications, winning global recognition for its training excellence, including 20 MCT Trainers in Microsoft’s Global Top 100 and an impressive 14 awards in the last 9 years. CloudThat specializes in Cloud Migration, Data Platforms, DevOps, Security, IoT, and advanced technologies like Gen AI & AI/ML. It has delivered over 750 consulting projects for 850+ organizations in 30+ countries as it continues to empower professionals and enterprises to thrive in the digital-first world.

FAQs

1. How many models can be compared?

ANS: – A single optimization job can compare the prompt across up to five inference models.

2. Is it the same as simple prompt optimization?

ANS: – No. Simple optimization performs a quick heuristic rewrite for one model. Advanced Prompt Optimization uses evaluation data, supports multiple models, and runs an iterative feedback process.

3. Does it require ground-truth answers?

ANS: – Ground-truth or reference responses are optional for some evaluation approaches. They are generally required when a custom deterministic metric needs to compare generated and expected outputs.

WRITTEN BY Yerraballi Suresh Kumar Reddy

Suresh is a highly skilled and results-driven Generative AI Engineer with over three years of experience and a proven track record in architecting, developing, and deploying end-to-end LLM-powered applications. His expertise covers the full project lifecycle, from foundational research and model fine-tuning to building scalable, production-grade RAG pipelines and enterprise-level GenAI platforms. Adept at leveraging state-of-the-art models, frameworks, and cloud technologies, Suresh specializes in creating innovative solutions to address complex business challenges.

Share

Comments

    Click to Comment

Get The Most Out Of Us

Our support doesn't end here. We have monthly newsletters, study guides, practice questions, and more to assist you in upgrading your cloud career. Subscribe to get them all!