Fei Xue Is Building Smarter and More Resilient Infrastructure for the AI Era

September 02 22:45 2026

Seattle, USA – September 2, 2026 – As AI becomes increasingly integral to software, the focus of software engineers is starting to change. While the development of better models is becoming more critical than ever before, so does the infrastructure supporting these models: the distributed systems, the cloud-based platforms, APIs, and edge devices all have to keep functioning as the systems become bigger and more complex.

For Fei Xue, the connection between AI and computing infrastructure is becoming a central theme not just in his engineering career but in his research endeavors as well. Rather than thinking about AI applications simply in terms of the end users, Xue has investigated how machine learning could be used in systems themselves to help them detect failures, to test their own resiliency, to allocate resources in smarter ways, and to react to changing environmental conditions.

Fei Xue currently works as a Software Development Engineer for Amazon while working towards his PhD in Information Technology. Previously, he obtained his Master of Science in Computer Engineering from the University of California, San Diego following his Bachelor of Engineering in Information Engineering from Shanghai Jiao Tong University. His background is focused entirely on backend engineering and large-scale business systems. In his career at Amazon, he has been working on platforms around invoicing, tax, payments, compliance, and procurement, such as PrintHub, which is a generalized platform capable of supporting many country-specific and business workflows. In addition, he has also worked on BOON – an AI-based internal project that aims to automate the process of customer onboarding and procurement.

This unique blend of large-scale engineering experience and AI research led Xue to pursue a particular research question that arises frequently from his practice: How can complex software systems become more adaptive to the environment in which they operate?

The question becomes particularly important when it comes to microservice architectures. Modern online systems can consist of tens or hundreds of interdependent services. While such architecture allows systems to scale quickly, it also means that any failure can propagate through a complex chain of dependencies between various services. Slowdown of one service can affect other services, lead to retried calls and cause failure that will eventually grow bigger and bigger.

This work from Xue attempts to solve the problem of fuzzing microservice resilience under uncertainty by introducing large language models to failure testing. Unlike traditional testing which involves a pre-defined set of fault cases, the paper studies the potential of LLM-enabled scenario generation to yield richer combinations of failures. The basic idea represents a general shift in the mindset of software reliability where engineers not only need to test if their system survives expected faults but also to find unforeseen combinations of events using intelligent tools.

However, detecting the failure is just one half of the issue at hand. In case of failure engineers have to understand its cause.

That is why CausalRCA, another work by Xue that focuses on causal graph-augmented retrieval for root cause analysis of microservices systems, has emerged. This research links AI reasoning with the structure and operation of the system. Unlike LLMs that act as prediction engines on their own, CausalRCA incorporates knowledge about service interrelations and retrieval-based causal reasoning into the process.

The importance of the relationship between elements also emerges in Xue’s study of topology-aware graph neural networks for distributed node fault identification. Distributed systems are, by nature, networked systems: Nodes in distributed systems interact with one another, rely on one another, and even pass on the results of faults through those interactions. Independent consideration of individual nodes may thus be insufficient.

Xue et al., instead, represent the distributed systems in a graph, thereby allowing the graph neural network to incorporate the features of individual nodes, along with information about neighboring nodes and system topology. Their experimental results include 0.912 accuracy, 0.938 AUC, and 0.891 F1 score; these results are superior to the performance of several other models. The researchers further analyze the effects of the model receiving information about more neighboring nodes, observing that while structural context is valuable for identification, too much information leads to redundancy.

But for Xue, intelligent infrastructure is not merely about identifying failures; it also involves making the right decisions in advance of any problem.

The approach of Xue in intelligent backend service scheduling applies reinforcement learning in conjunction with meta-learning to overcome the limitations imposed by a fixed resource allocation policy. While traditional scheduling algorithms use predefined thresholds and policies which are useful in known environments, they become increasingly ineffective in situations when workloads, available resources or service configurations change.

The scheduling problem is modelled by Xue as a sequence of decisions and reinforcement learning enables the algorithm to optimize the policy used in resource allocations and meta-learning is utilized to adapt the resource allocation policy faster in response to new workloads. With Alibaba Cluster Trace 2018 dataset, the suggested approach achieved the average latency of 143.7 ms, CPU utilization of 83.6%, and the SLO violation rate of 2.7% during comparative experiments performed.

In addition, some other topics addressed in Xue’s research papers apply the same approach in different areas. For example, the work on requirement-aware LLM testing for backend REST APIs explores the issue of testing of AI-generated software tests against the requirements that the software satisfies. The problem is essential as technically correct tests and high code coverage do not guarantee that the software fulfills the business logic. By tying the requirements to API testing, the study shows the potential of generative AI being more grounded in real needs of software engineering.

On the other hand, his work on efficient distributed inference in the edge-cloud continuum tackles another limitation: AI models require computing and energy that are quite diverseamong devices. Model compression is an approach to adapt the size and complexity of the models depending on the availability of the resources in order to enable intelligent apps to function more effectively beyond centralized cloud computing.

All taken together, these studies show a research agenda that goes far beyond any particular method of machine learning. Graph neural networks, large language models, reinforcement learning, meta-learning, causal reasoning, and model compression are just some examples from Xue’s research agenda, yet the underlying goal remains the same.

He is looking at ways to bring intelligence into infrastructure itself.

This is because the need for systems that can introspect their vulnerabilities before deployment, analyze failure causes in case of any, comprehend the interactions between distributed elements, allocate resources according to changes in the environment, prove that software meets all requirements and modify AI computations based on hardware limitations is needed.

With the increasing complexity of software systems and the prevalence of AI technologies in computing, this could be the case. It could be that the next phase of intelligent computing requires not just better algorithms but also a system infrastructure that can support them.

Xue’s research is an indicator of that future, where computing systems will not only perform instructions but learn to monitor, analyze, adapt and optimize themselves.

About Fei Xue

Fei Xue is a software engineer and researcher specialized in developing reliable and intelligent computing systems. His research areas include artificial intelligence, distributed systems, cloud computing infrastructure, reliability of microservices, and software engineering with artificial intelligence. Xue works as a Software Development Engineer II for Amazon while studying for his Doctorate of Information Technology. He holds an M.S. in Computer Engineering from the University of California, San Diego, and a B.S. in Information Engineering from Shanghai Jiao Tong University.

Media Contact
Company Name: Fei Xue
Contact Person: Media Relations
Email: Send Email
City: Seattle
Country: United States
Website: https://www.linkedin.com/in/feixue94b641193/?locale=en