کتاب Learning AutoML راهنمای عملی شما برای بهکارگیری یادگیری ماشین خودکار (AutoML) در محیطهای واقعی است. این کتاب به شما کمک میکند تا از مرحله آزمایش فراتر رفته و مدلهایی با عملکرد بالا را با اتوماسیون بیشتر و تنظیمات دستی کمتر بسازید و مستقر کنید. با استفاده از AutoGluon به عنوان ابزار اصلی، یاد میگیرید که چگونه مدلهای AutoML را بسازید، ارزیابی کنید و به کار بگیرید تا پیچیدگی را کاهش داده و نوآوری را تسریع کنید. نویسنده، Kerem Tomak، بینشهایی را در مورد نحوه ادغام مدلها در جریانهای کاری استقرار سرتاسری با استفاده از ابزارهای محبوبی مانند Kubeflow، MLflow و Airflow به اشتراک میگذارد.
اشتراکی

با مرور فصلها، ساختار ، محتوای کتاب را به سرعت بشناسید.
با مرور فصلهای این کتاب میتونی خیلی سریع بفهمی هر بخش چی یاد میده، ساختار کلی چطوره و از کجا باید شروع کنی. هر فصل روی یک مفهوم یا مهارت خاص تمرکز داره و موضوعات اصلیش رو میبینی تا انتخابت آگاهانهتر باشه. چه بخوای کل کتاب رو دنبال کنی، چه فقط یک بخش خاص رو دنبال کنی، این نما کمکت میکنه مسیرت رو پیدا کنی.
در این بخش، ما با یک سوال اساسی شروع میکنیم: AutoML دقیقاً چیست و چرا برای متخصصان مدرن یادگیری ماشین ضروری شده است؟ ما تکامل آن را از تحقیقات دانشگاهی تا پلتفرمهای آماده برای شرکتها دنبال میکنیم.
مهمتر از همه، ما خط لوله سرتاسری AutoML را کالبدشکافی خواهیم کرد؛ ارکستراسیون پیچیدهای از پیشپردازش داده، مهندسی ویژگی، انتخاب مدل، بهینهسازی ابرپارامتر و ساخت گروهی که بهطور خودکار اتفاق میافتد.

یادگیری ماشین خودکار چیست؟
The Growing Demand for Machine Learning Solutions • Addressing the Data Science Talent Gap • Democratizing AI Development • AutoML in the Machine Learning Landscape • Open Source AutoML Libraries • Enterprise AutoML Platforms • Comparison of Leading Frameworks • Who Should Use AutoML? • AutoML Across Industries: Transforming Business Processes • Finance • Healthcare and Life Sciences • Retail and Ecommerce • Manufacturing • Other Sectors • The Tiered Use Case Model • Overcoming Hurdles: Persistent Challenges in AutoML • Interpretability (the “Black Box” Problem) • Need for Customization Versus Automation • Data Quality Dependency and Robustness • Computational Costs and Resource Intensity • Addressing Bias and Fairness • Scalability and Efficiency • The Horizon: Future Trends Shaping AutoML • Synergy with Large Language Models (LLMs) and Foundation Models • Next-Generation Neural Architecture Search (NAS) • Maturity of Multimodal Explainable AI (MXAI) • Continued Democratization via Low-Code/No-Code • Expansion to Edge Computing and Federated Learning • Summary
درحال تولید...

ظهور و وضعیت فعلی AutoML
Early Automation (Pre-2010): Laying the Groundwork • Feature Selection • Hyperparameter Search • Meta-Learning Research • Limitations of Early Approaches • First Generation (2010–2015): Solving the CASH Problem • Auto-WEKA (2013) • Hyperopt (2013) • SMAC (Sequential Model-Based Algorithm Configuration) • First Generation Impact and Legacy • Second Generation (2015–2020): Solving the Usability and Enterprise Problem • Auto-sklearn (2015) • PyCaret (2020) • TPOT (Tree-Based Pipeline Optimization Tool) (2016) • H2O AutoML • Google Cloud AutoML (Now Part of Vertex AI) • Second Generation Impact and Legacy • Third Generation (2020–Present): Solving the Multimodal and MLOps Problem • AutoGluon (Amazon) • Google Vertex AI • MLJAR and AWS SageMaker Autopilot • Third Generation Key Capabilities • The Emergence of LLM-Assisted AutoML • Summary
درحال تولید...

درک خط لوله AutoML
The Architecture of Automated Machine Learning • Data Preprocessing • Data Quality Assessment and Cleaning • Missing Data Strategies • Data Validation and Integrity Checks • Feature Engineering • Multilevel Feature Generation • Domain-Specific Feature Engineering • Feature Selection and Pruning • Representation Learning Integration • Hyperparameter Optimization • Advanced Search Strategies • Multifidelity Optimization • Configuration Space Design • Budget-Aware Optimization • Neural Architecture Search • Search Space Engineering • Efficiency Techniques • Hardware-Aware Architecture Search • Architecture Transfer and Meta-Learning • Model Selection, Ensembling, and Stacking • Diversity-Driven Ensemble Construction • Advanced Stacking Techniques • Dynamic Ensemble Selection • Resource-Aware Ensemble Optimization • Model Deployment and Monitoring • Production Readiness Considerations • Scalability and Performance Optimization • Model Monitoring and Maintenance • Interpretability and Explainability • Pipeline Integration and Optimization • Cross-Stage Optimization Strategies • Resource Allocation and Management • Feedback Mechanisms and Continuous Learning • Challenges and Future Directions • Scalability and Efficiency • Robustness and Reliability • Democratization and Accessibility • The Double-Edged Sword of Democratization • Summary
درحال تولید...
جادوی AutoML در سه تکنیک اصلی نهفته است که طی دههها تحقیق اصلاح شدهاند: مهندسی ویژگی خودکار، بهینهسازی ابرپارامتر و جستجوی معماری عصبی. درک عمیق این تکنیکها به شما کمک میکند تا تصمیمات بهتری بگیرید.
بهینهسازی ابرپارامتر و جستجوی معماری عصبی، قلب الگوریتمی AutoML را تشکیل میدهند. ما به بررسی بهینهسازی بیزی، الگوریتمهای تکاملی و استراتژیهای جستجویی میپردازیم که به سیستمهای AutoML اجازه میدهند تا فضاهای پیکربندی وسیع را بهطور کارآمد مدیریت کنند.

پیشپردازش خودکار داده و مهندسی ویژگی
Working Dataset: RetailMart Ecommerce Platform • Intelligent Data Profiling and Quality Assessment • Smart Data Type Handling and Transformation • DateTime Feature Extraction • Text Preprocessing Pipelines • Automated Feature Engineering • Traditional Feature Engineering Automation • Advanced Feature Learning Techniques • Intelligent Feature Selection and Dimensionality Management • Preprocessing Complex and Multimodal Data • Production-Ready Preprocessing Pipelines • Summary
درحال تولید...

بهینهسازی ابرپارامتر
The Challenge of Hyperparameter Optimization • The Computational Cost Challenge • The Sensitivity Problem • Real-World Impact • Grid Search Versus Random Search: Building the Foundation • Grid Search: Systematic but Limited • Random Search: A Surprisingly Effective Alternative • A Practical Comparison • Modern Implementation • When to Use Each Approach • Limitations of Both Approaches • Bayesian Optimization: Learning from Experience • The Core Insight • Surrogate Models • Acquisition Functions • Real-World Success Stories • Modern Tools and Implementation • AWS SageMaker Automatic Model Tuning • Advanced Techniques • Practical Considerations • Limitations and Challenges • Early Stopping and Scheduling: Working Smarter, Not Harder • The Core Insight • Successive Halving: A Tournament Approach • Hyperband: Automating the Resource Allocation • Asynchronous Successive Halving (ASHA) • Population-Based Training: Evolution During Training • Layer Freezing: A Novel Fidelity Dimension • Practical Implementation • Real-World Results • Combining with Bayesian Optimization • When Early Stopping Works Best • A Cautionary Note • Multifidelity Optimization: Beyond Simple Early Stopping • The Multifidelity Paradigm • Advanced Multifidelity Methods • Practical Implementation • Case Study: Personal Investment Portfolio Optimization with Multifidelity HPO • Background and Problem Definition • Dataset and Features • Multifidelity Strategy Implementation • Implementation • Resource Management and Results • Key Insights and Practical Considerations • Early Stopping Effectiveness • Model Type Performance Patterns • Resource Allocation Strategy • Production Deployment Considerations • Lessons for Individual Practitioners • When to Use Multifidelity Optimization • Summary
درحال تولید...

جستجوی معماری عصبی (NAS)
Understanding Neural Architecture Search • The Three Pillars of NAS • Search Space Design: Defining the Boundaries • The Art of Constraint • Types of Search Spaces • Task-Specific Considerations • Emerging Specialized Search Spaces • Multi-Objective Search Spaces • Balancing Efficiency and Discovery • The NAS-Bench Revolution • Search Strategies: Finding Needles in Haystacks • The Evolution of Search Strategies • Choosing Your Search Strategy: A Practical Perspective • Reinforcement Learning: The Original Approach • Evolutionary Algorithms: Nature-Inspired Search • Differentiable NAS: The Game Changer • Gradient-Based Methods and Advanced Techniques • Hybrid Approaches: Best of Both Worlds • Choosing the Right Strategy • Performance Estimation: The Efficiency Imperative • Bridging the Evaluation Gap • The Training Bottleneck • Multifidelity Evaluation: Training Less, Learning More • One-Shot Architecture Search: Train Once, Evaluate Many • Learning Curve Extrapolation • Zero-Cost Proxies: Instant Architecture Evaluation • Surrogate Models: Learning to Predict Performance • Combining Approaches for Maximum Efficiency • Efficient NAS: Making It Practical • The Efficiency Revolution • Building Production-Ready NAS Systems • Weight Sharing: The Foundation of Efficient NAS • Once-For-All Networks: Decoupling Training and Deployment • Progressive Search Strategies • Hardware-Aware Optimization • Production Deployment Tools for NAS-Discovered Architectures • Zero-Cost Proxies for Rapid Filtering • Practical Implementation Guidelines • Practical Applications and Tools • NAS in the Real World: Integration and Deployment • AutoKeras: Simplicity First • NNI: Enterprise-Grade NAS • Ray Tune + Optuna: Flexible and Powerful • Industry Success Stories • From Notebook to Production: Next Steps • Summary
درحال تولید...
در این بخش، ما به پیادهسازی عملی در چهار نوع داده اصلی که در تولید با آنها مواجه خواهید شد، میپردازیم. برای دادههای جدولی، بررسی میکنیم که چرا مدلهای تقویت گرادیان (Gradient Boosting) همچنان غالب هستند.
برای متن، یاد میگیرید که چگونه از مدلهای مبتنی بر ترنسفورمر استفاده کنید. پیشبینی سریهای زمانی، وابستگیهای زمانی را معرفی میکند و بینایی ماشین نشان میدهد که چگونه AutoML قابلیتهایی را که زمانی به تیمهای تخصصی نیاز داشت، دموکراتیزه میکند.

AutoGluon برای دادههای جدولی
Setting Up AutoGluon and Environment • Installation Options • Platform-Specific Notes • Setting Up Your Development Environment • Performance Considerations • Cloud Environment Recommendations • Choosing the Right AutoML Framework for Tabular Data • TabularPredictor Basics • Loading and Exploring Data • Basic Model Training • Understanding TabularPredictor Output • Different Prediction Methods • Binary and Multiclass Classification • Binary Classification in Detail • Multiclass Classification • Regression Tasks • Regression Versus Classification Differences • Interpreting Regression Performance • Customizing Basic Behavior • AutoGluon’s Automatic Data Processing • Automatic Feature Type Detection • Missing Value Handling • Categorical Encoding • Advanced Customization • Custom Hyperparameters • Advanced Ensemble Configuration • Feature Engineering Control • Training Process Optimization • Model Interpretability and Debugging • Interpretability Tools • Handling Special Data Scenarios • When to Use Advanced Customization • Project: Titanic Survival Prediction • Project Overview and Business Context • Data Exploration and Understanding • Baseline AutoGluon Model • Custom Feature Engineering for Titanic • Model Interpretation for Titanic • Performance Evaluation and Comparison • Model Deployment Preparation • Project Summary and Business Impact • Extending This Project • Data Pipeline Consistency • Monitoring and Maintaining Models in Production • Monitoring Practices • Monitoring Tools for AutoGluon • Summary
درحال تولید...

AutoML برای متن و پردازش زبان طبیعی
AutoGluon’s MultiModalPredictor for Text Processing • Why MultiModalPredictor? • Underlying Model Architectures • Real-World Performance • Building Text Classification Models • Your First Text Classification Model • Understanding Model Selection • Hyperparameter Optimization Guidelines • Advanced Text Processing Capabilities • Beyond Classification: Advanced NLP Tasks • The Transformer Revolution and Beyond • Domain-Specific Considerations • Model Selection for Different Use Cases • Maximum Accuracy Applications • Balanced Applications • High-Throughput Applications • Real-World Applications and Performance • Industry Case Studies • Performance Insights • Production Deployment Considerations • Model Selection for Deployment Scenarios • Managed Services Versus Custom Models • Deploying Custom Models with SageMaker • Monitoring and Maintenance • Performance and Operational Monitoring • Data Drift Detection • Retraining and Continuous Improvement • Practical Project: News Article Classification • Summary
درحال تولید...

پیشبینی سریهای زمانی با AutoGluon
Understanding the Time Series Challenge • Getting Started with TimeSeriesPredictor • Foundation Models and Zero-Shot Forecasting • The Chronos-Bolt Architecture • Real-World Impact of Zero-Shot Forecasting • Handling Complex Multiseries Scenarios • Advanced Capabilities: Covariate Regressors • Implementing Covariate Regressors • Business Impact of Covariate Integration • Model Selection and Hyperparameter Optimization • The Model Zoo • Preset Configurations • Custom Hyperparameter Configuration • Evaluation and Validation Strategies • Backtesting and Time-Aware Validation • Business-Relevant Metrics • Production Deployment and Cloud Integration • AWS Deployment Options • Model Updating and Monitoring • Practical Project: Retail Demand Forecasting • Data Preparation and Exploration • Model Training with Advanced Features • Business Impact Analysis • Future Directions and Emerging Capabilities • Summary
درحال تولید...

بینایی ماشین با AutoGluon
Understanding AutoGluon’s Computer Vision Capabilities • Choosing Between Custom Models and Managed Services • Building Training Datasets with SageMaker Ground Truth • The MultiModalPredictor Advantage • Foundation Models Integration • Modern Computer Vision Architectures • Task Categories and Applications • Setting Up AutoGluon for Computer Vision • Installation and Environment Setup • Hardware Considerations • Verification and Basic Setup • Image Classification with MultiModalPredictor • Your First Image Classification Model • Understanding Data Formats and Preprocessing • Model Architecture Selection and Presets • Advanced Classification Techniques • Object Detection with AutoGluon • Understanding Object Detection • Basic Object Detection Setup • Enhanced Object Detection Capabilities • Advanced Object Detection Applications • Multimodal Computer Vision Applications • Combining Images with Tabular Data • Image and Text Integration • Real-World Computer Vision Project: Automated Ecommerce Product Classification • Project: Automated Ecommerce Product Classification • Data Preparation and Exploration • Building the Multimodal Classification System • Performance Analysis and Model Interpretability • Integration with Ecommerce Systems • Performance Optimization and Best Practices • Hardware Optimization Strategies • Model Monitoring and Maintenance • Production Deployment Considerations • Model Versioning and Updates • SageMaker Endpoint Deployment • SageMaker Serverless Inference for Cost-Effective Deployment • AWS Panorama for Edge Deployment • Scalable Batch Processing Service • Summary
درحال تولید...
ساخت مدلهای دقیق تنها نیمی از راه است. نیمه دیگر—که اغلب سختتر است—رساندن این مدلها به تولید، حفظ عملکرد قابل اطمینان آنها و اطمینان از عملکرد صحیح آنها در حین تغییر جهان است.
در این بخش، ما شکاف بین AutoML و MLOps را پر میکنیم. شما یاد میگیرید که AutoML را در پلتفرمهای مدرن ML مانند MLflow ادغام کنید، خط لولههای داده قابل اطمینان با Apache Airflow بسازید و شیوههای CI/CD را پیادهسازی کنید که با استقرار مدل با همان دقتی برخورد میکنند که با استقرار نرمافزار.

یکپارچهسازی جریان کاری با ابزارهای MLOps
Understanding the AutoML-MLOps Integration Landscape • The Scale Challenge • The Reproducibility Imperative • Experiment Tracking and Model Management • Hierarchical Experiment Organization • Artifact Management Strategies • Workflow Orchestration with Kubeflow • Designing AutoML-Aware Pipelines • Resource Management and Optimization • Production Deployment Patterns • Automated Validation and Quality Assurance • Dynamic Serving Infrastructure • Operational Monitoring and Maintenance • Monitoring and Governance • AutoML-Specific Monitoring Requirements • Governance and Compliance Frameworks • Integration Challenges and Solutions • The Artifact Explosion Challenge • Ensuring Reproducibility in Automated Systems • Bridging Technical and Business Domains • Best Practices and Implementation Guidelines • Building Progressive Capability • Alignment and Expectation Management • Risk Management and Parallel Systems • Observability as a Foundation • Organizational Learning and Adaptation • Summary
درحال تولید...

اتوماسیون خط لوله داده با Apache Airflow
Understanding Data Pipeline Requirements for AutoML • Airflow Architecture for Machine Learning Workflows • Core Components • Key Airflow Terminology • Designing DAGs for AutoML Data Ingestion • Practical Example: Complete AutoML Data Ingestion DAG • DAG Initialization and Configuration • Understanding catchup Behavior • Dynamic Task Mapping for Parallel Processing • Feature Engineering Pipelines and Feature Stores • Handling Late-Arriving Data • Data Contracts and Schema Evolution • Monitoring and Data Quality Gates • Scaling Airflow for Enterprise AutoML • Operational Excellence and Best Practices • Summary
درحال تولید...

استقرار و تحویل مستمر برای AutoML
The Unique Challenges of AutoML Deployment • Continuous Integration for Machine Learning • Shadow Deployment Validation • Continuous Deployment Pipelines • Testing Strategies for Automated Models • Contract Testing • Property-Based Testing • Metamorphic Testing • Adversarial Testing • Model Packaging and Containerization • Practical Example: Deploying the Adult Income Prediction Model • Model Serving Infrastructure • Monitoring and Observability in Production • Prometheus–Grafana Monitoring Stack • Drift Detection with Evidently • Security and Compliance Considerations • Input Sanitization and DoS Prevention • Adversarial Attack Defenses • Continuous Learning and Feedback Loops • Summary
درحال تولید...
در این بخش، ما به بررسی کاربردهای عملی AutoML در صنایع مختلف میپردازیم. از تشخیص تقلب در خدمات مالی گرفته تا پیشبینی تقاضا در خردهفروشی و پیشبینی بازگشت بیمار در مراقبتهای بهداشتی، این مطالعات موردی نشان میدهند که چگونه اتوماسیون میتواند پروژههایی را که ماهها زمان میبردند، به روزها یا حتی ساعتها کاهش دهد.

مطالعه موردی ۱: خدمات مالی—تشخیص تقلب در زمان واقعی در GlobalBank
Business Problem and Context • Success Criteria • Data Pipeline and Preparation • Data Pipeline Architecture • Production Data Pipeline Considerations • Feature Engineering • 1. Temporal Features: Fraud Has a Schedule • 2. Velocity Features: Fraudsters Move Fast • 3. Behavioral Deviation: Detecting the Unusual • 4. Merchant Risk Scoring • 5. Device Trust • Feature Impact Summary • Model Development with AutoGluon • Sample Weighting for Cost-Sensitive Learning • AutoGluon Configuration • Why PR-AUC Instead of ROC-AUC? • Model Training Results • Model Evaluation and Interpretability • Finding the Optimal Threshold • Model Interpretability with SHAP • Deployment Architecture • FastAPI Inference Service • Graceful Degradation Strategy • Monitoring and Maintenance • Drift Detection with PSI • Automated Retraining Pipeline • A/B Testing for Model Updates • Outcomes and Lessons Learned • Performance Metrics • Business Impact • Key Lessons • The Right Division of Labor • Summary
درحال تولید...

مطالعه موردی ۲: خردهفروشی—پیشبینی تقاضای همهکاناله
Business Problem and Context • The Scale Challenge • The Wake-Up Call • Project Objectives • Data Challenges: Multisource Integration • Point-of-Sale Data • Ecommerce Data • Inventory Data • Marketing and Promotions Data • Weather Data • External Signals • Data Pipeline Architecture • Key Data Decisions • Feature Engineering: Capturing Demand Drivers • Business-Aligned Metrics • Temporal Features (Baseline) • Omnichannel Behavioral Features (High Impact) • Weather-Driven Demand (Category-Specific) • Promotional Features (Complex Interactions) • Event-Driven Demand • SKU-Specific Attributes • Model Development: AutoGluon for Time Series at Scale • The AutoML Approach • Why Tabular AutoML for Time Series? • Training Strategy: Time-Based Splits • Multihorizon Forecasting • AutoGluon Configuration • Key Configuration Decisions • Handling Data Sparsity (Long-Tail SKUs) • Training Infrastructure • Evaluation: Business Metrics over Model Metrics • Model Performance (MAPE by Forecast Horizon) • Weighted Average MAPE (Business-Aligned) • Business Impact Metrics • Category-Specific Performance • Promotional Forecast Accuracy • Forecast Bias Analysis • Deployment: Production Forecasting Pipeline • Pipeline Architecture • Technology Stack • Forecast Serving • Monitoring: Keeping Forecasts Accurate • Drift Detection • Business Outcomes and Lessons Learned • Quantified Business Impact (12 Months Postlaunch) • Unexpected Benefits • Critical Success Factors • What We’d Do Differently • Lessons for Your Demand Forecasting Project • When AutoML Excels for Demand Forecasting • Summary
درحال تولید...

مطالعه موردی ۳: مراقبتهای بهداشتی—پیشبینی بازگشت بیمار
The Business Challenge • The Hardest Constraint: Fairness • The Current State • Project Objectives • Data Challenges and HIPAA Compliance • Data Sources and Integration • Missing Data Patterns • Data Quality Issues • Feature Engineering: Structured and Unstructured Data • Category 1: Demographics and Social Determinants (42 Features) • Category 2: Clinical Complexity and Comorbidities (68 features) • Category 3: Utilization History (53 features) • Category 4: Current Encounter Features (87 features) • Category 5: Clinical Notes Embeddings (64 features) • Category 6: Temporal and Interaction Features (33 features) • Model Development: Fairness-Aware AutoML • The Fairness Challenge • Fairness Metrics Defined • Baseline Model: Standard AutoGluon (No Fairness Constraints) • Approach 1: Remove Protected Attributes • Approach 2: Adversarial Debiasing • Approach 3 (Final): Fairness-Aware Ensemble with Reweighting • Final Model Configuration • Evaluation: Performance + Fairness Metrics • Model Performance (Overall) • Fairness Metrics by Race • Fairness by Age Group • Feature Importance (Top 20 by SHAP) • Business Metrics • Deployment: Clinical Workflow Integration • Real-Time Prediction Architecture • EHR Integration (Epic) • Clinical Decision Support Alert • Care Management Workflow • Clinician Training and Change Management • Interpretability for Clinicians • Regulatory Considerations • Monitoring: Drift and Fairness • Three-Layer Monitoring Strategy • Retraining Schedule • Business Outcomes and Lessons Learned • Clinical Outcomes • Fairness in Practice • Unexpected Benefits • Critical Success Factors • What We’d Do Differently • Lessons for Your Readmission Project • When AutoML Excels for Healthcare • The Production AutoML Blueprint: A Grand Synthesis • The Universal Patterns of Production AutoML • The Production Readiness Checklist • Final Thoughts • Summary
درحال تولید...
16 فصل در حال تولید
مدت زمان خوانش
16:26
نوع کتاب
اشتراکی
شرکت کنندگان
0 نفر
تولید کتاب
۳۱ شهریور ۱۴۰۵