关于云原生:使用-Arthas-排查-SpringBoot-诡异耗时的-Bug

47次阅读

共计 31982 个字符,预计需要花费 80 分钟才能阅读完成。

作者 | 空无
起源 | 阿里巴巴云原生公众号

背景

公司有个渠道零碎,专门对接三方渠道应用,没有什么业务逻辑,次要是转换报文和参数校验之类的工作,起着一个承前启后的作用。

最近在优化接口的响应工夫,优化了代码之后,然而工夫还是达不到要求;有一个诡异的 100ms 左右的耗时问题,在接口中打印了申请解决工夫后,和调用方的响应工夫还有差了 100ms 左右。比方程序里记录 150ms,然而调用方等待时间却为 250ms 左右。

上面记录下过后具体的定位 & 解决流程(其实解决很简略,关键在于怎么定位并找到解决问题的办法)。

定位过程

1. 剖析代码

渠道零碎是一个常见的 Spring-boot web 工程,应用了集成的 tomcat。剖析了代码之后,发现并没有非凡的中央,没有非凡的过滤器或者拦截器,所以初步排除是业务代码问题。

2. 剖析调用流程

呈现这个问题之后,首先确认了下接口的调用流程。因为是内部测试,所以调用流程较少。

Nginx - 反向代理 -> 渠道零碎

公司是云服务器,网络走的也是云的内网。因为不明确问题的起因,所以用排除法,首先确认服务器网络是否有问题。

先确认发送端到 Nginx Host 是否有问题:

[jboss@VM_0_139_centos ~]$ ping 10.0.0.139
PING 10.0.0.139 (10.0.0.139) 56(84) bytes of data.
64 bytes from 10.0.0.139: icmp_seq=1 ttl=64 time=0.029 ms
64 bytes from 10.0.0.139: icmp_seq=2 ttl=64 time=0.041 ms
64 bytes from 10.0.0.139: icmp_seq=3 ttl=64 time=0.040 ms
64 bytes from 10.0.0.139: icmp_seq=4 ttl=64 time=0.040 ms

从 ping 后果上看,发送端到 Nginx 主机的提早是无问题的,接下来查看 Nginx 到渠道零碎的网络。

# 因为日志是没问题的,这里间接复制下面日志了
[jboss@VM_0_139_centos ~]$ ping 10.0.0.139
PING 10.0.0.139 (10.0.0.139) 56(84) bytes of data.
64 bytes from 10.0.0.139: icmp_seq=1 ttl=64 time=0.029 ms
64 bytes from 10.0.0.139: icmp_seq=2 ttl=64 time=0.041 ms
64 bytes from 10.0.0.139: icmp_seq=3 ttl=64 time=0.040 ms
64 bytes from 10.0.0.139: icmp_seq=4 ttl=64 time=0.040 ms

从 ping 后果上看,Nginx 到渠道零碎服务器网络提早也是没问题的。

既然网络看似没问题,那么能够持续排除法,砍掉 Nginx,客户端间接再渠道零碎的服务器上,通过回环地址(localhost)直连,防止通过网卡 /dns,放大问题范畴看看是否复现(这个利用和地址是我前期模仿的,测试的是一个空接口):

[jboss@VM_10_91_centos tmp]$ curl -w "@curl-time.txt" http://127.0.0.1:7744/send
success
              http: 200
               dns: 0.001s
          redirect: 0.000s
      time_connect: 0.001s
   time_appconnect: 0.000s
  time_pretransfer: 0.001s
time_starttransfer: 0.073s
     size_download: 7bytes
    speed_download: 95.000B/s
                  ----------
        time_total: 0.073s 申请总耗时

从 curl 日志上看,通过回环地址调用一个空接口耗时也有 73ms。这就奇怪了,跳过了两头所有调用节点(包含过滤器 & 拦截器之类),间接申请利用一个空接口,都有 73ms 的耗时,再申请一次看看:

[jboss@VM_10_91_centos tmp]$ curl -w "@curl-time.txt" http://127.0.0.1:7744/send
success
              http: 200
               dns: 0.001s
          redirect: 0.000s
      time_connect: 0.001s
   time_appconnect: 0.000s
  time_pretransfer: 0.001s
time_starttransfer: 0.003s
     size_download: 7bytes
    speed_download: 2611.000B/s
                  ----------
        time_total: 0.003s

更奇怪的是,第二次申请耗时就失常了,变成了 3ms。经查阅材料,linux curl 是默认开启 http keep-alive 的(Keep-Alive 的介绍能够参考我的另一篇文章)。就算不开启 keep-alive,每次从新 handshake,也不至于须要 70ms。

通过一直分析测试发现,间断申请的话工夫就会很短,每次申请只须要几毫秒,然而如果隔一段时间再申请,就会破费 70ms 以上。

从这个景象猜测,可能是某些缓存机制导致的,间断申请因为有缓存,所以速度快,工夫长缓存生效后导致工夫长。

那么这个问题点到底在哪一层呢?tomcat 层还是 spring-webmvc 呢?

光猜测定位不了问题,还是得理论测试一下,把渠道零碎的代码放到本地 IDE 里启动测试是否复现。

然而导入本地 IDE 后,在 IDE 中启动后并不能复现问题,并没有 70+ms 的提早问题。这下头疼了,本地无奈复现,不能 Debug,因为问题点不在业务代码,也不能通过加日志的形式来 Debug。

这时候能够祭出神器 Arthas 了

3. Arthas 剖析问题

Arthas 是 Alibaba 开源的 Java 诊断工具,深受开发者青睐。当你遇到以下相似问题而大刀阔斧时,Arthas 能够帮忙你解决:

  • 这个类从哪个 jar 包加载的?为什么会报各种类相干的 Exception?
  • 我改的代码为什么没有执行到?难道是我没 commit?分支搞错了?
  • 遇到问题无奈在线上 debug,难道只能通过加日志再从新公布吗?
  • 线上遇到某个用户的数据处理有问题,但线上同样无奈 debug,线下无奈重现!
  • 是否有一个全局视角来查看零碎的运行状况?
  • 有什么方法能够监控到 JVM 的实时运行状态?
  • ······

下面是 Arthas 的官网简介,这次我只须要用他的一个小性能 trace。动静计算方法调用门路和工夫,这样就能够定位工夫在哪个中央被耗费了。

trace 办法外部调用门路,并输入办法门路上的每个节点上耗时。

trace 命令能被动搜寻 class-pattern/method-pattern。

对应的办法调用门路,渲染和统计整个调用链路上的所有性能开销和追踪调用链路。

有了神器,那么该追踪什么办法呢?因为我对 Tomcat 源码不是很熟,所以只能从 spring mvc 下手,先来 trace 一下 spring mvc 的入口:

[arthas@24851]$ trace org.springframework.web.servlet.DispatcherServlet *
Press Q or Ctrl+C to abort.
Affect(class-cnt:1 , method-cnt:44) cost in 508 ms.
`---ts=2019-09-14 21:07:44;thread_name=http-nio-7744-exec-2;id=11;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917
    `---[2.952142ms] org.springframework.web.servlet.DispatcherServlet:buildLocaleContext()
`---ts=2019-09-14 21:07:44;thread_name=http-nio-7744-exec-2;id=11;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917
    `---[18.08903ms] org.springframework.web.servlet.DispatcherServlet:doService()
        +---[0.041346ms] org.apache.commons.logging.Log:isDebugEnabled() #889
        +---[0.022398ms] org.springframework.web.util.WebUtils:isIncludeRequest() #898
        +---[0.014904ms] org.springframework.web.servlet.DispatcherServlet:getWebApplicationContext() #910
        +---[1.071879ms] javax.servlet.http.HttpServletRequest:setAttribute() #910
        +---[0.020977ms] javax.servlet.http.HttpServletRequest:setAttribute() #911
        +---[0.017073ms] javax.servlet.http.HttpServletRequest:setAttribute() #912
        +---[0.218277ms] org.springframework.web.servlet.DispatcherServlet:getThemeSource() #913
        |   `---[0.137568ms] org.springframework.web.servlet.DispatcherServlet:getThemeSource()
        |       `---[min=0.00783ms,max=0.014251ms,total=0.022081ms,count=2] org.springframework.web.servlet.DispatcherServlet:getWebApplicationContext() #782
        +---[0.019363ms] javax.servlet.http.HttpServletRequest:setAttribute() #913
        +---[0.070694ms] org.springframework.web.servlet.FlashMapManager:retrieveAndUpdate() #916
        +---[0.01839ms] org.springframework.web.servlet.FlashMap:<init>() #920
        +---[0.016943ms] javax.servlet.http.HttpServletRequest:setAttribute() #920
        +---[0.015268ms] javax.servlet.http.HttpServletRequest:setAttribute() #921
        +---[15.050124ms] org.springframework.web.servlet.DispatcherServlet:doDispatch() #925
        |   `---[14.943477ms] org.springframework.web.servlet.DispatcherServlet:doDispatch()
        |       +---[0.019135ms] org.springframework.web.context.request.async.WebAsyncUtils:getAsyncManager() #953
        |       +---[2.108373ms] org.springframework.web.servlet.DispatcherServlet:checkMultipart() #960
        |       |   `---[2.004436ms] org.springframework.web.servlet.DispatcherServlet:checkMultipart()
        |       |       `---[1.890845ms] org.springframework.web.multipart.MultipartResolver:isMultipart() #1117
        |       +---[2.054361ms] org.springframework.web.servlet.DispatcherServlet:getHandler() #964
        |       |   `---[1.961963ms] org.springframework.web.servlet.DispatcherServlet:getHandler()
        |       |       +---[0.02051ms] java.util.List:iterator() #1183
        |       |       +---[min=0.003805ms,max=0.009641ms,total=0.013446ms,count=2] java.util.Iterator:hasNext() #1183
        |       |       +---[min=0.003181ms,max=0.009751ms,total=0.012932ms,count=2] java.util.Iterator:next() #1183
        |       |       +---[min=0.005841ms,max=0.015308ms,total=0.021149ms,count=2] org.apache.commons.logging.Log:isTraceEnabled() #1184
        |       |       `---[min=0.474739ms,max=1.19145ms,total=1.666189ms,count=2] org.springframework.web.servlet.HandlerMapping:getHandler() #1188
        |       +---[0.013071ms] org.springframework.web.servlet.HandlerExecutionChain:getHandler() #971
        |       +---[0.372236ms] org.springframework.web.servlet.DispatcherServlet:getHandlerAdapter() #971
        |       |   `---[0.280073ms] org.springframework.web.servlet.DispatcherServlet:getHandlerAdapter()
        |       |       +---[0.004804ms] java.util.List:iterator() #1224
        |       |       +---[0.003668ms] java.util.Iterator:hasNext() #1224
        |       |       +---[0.003038ms] java.util.Iterator:next() #1224
        |       |       +---[0.006451ms] org.apache.commons.logging.Log:isTraceEnabled() #1225
        |       |       `---[0.012683ms] org.springframework.web.servlet.HandlerAdapter:supports() #1228
        |       +---[0.012848ms] javax.servlet.http.HttpServletRequest:getMethod() #974
        |       +---[0.013132ms] java.lang.String:equals() #975
        |       +---[0.003025ms] org.springframework.web.servlet.HandlerExecutionChain:getHandler() #977
        |       +---[0.008095ms] org.springframework.web.servlet.HandlerAdapter:getLastModified() #977
        |       +---[0.006596ms] org.apache.commons.logging.Log:isDebugEnabled() #978
        |       +---[0.018024ms] org.springframework.web.context.request.ServletWebRequest:<init>() #981
        |       +---[0.017869ms] org.springframework.web.context.request.ServletWebRequest:checkNotModified() #981
        |       +---[0.038542ms] org.springframework.web.servlet.HandlerExecutionChain:applyPreHandle() #986
        |       +---[0.00431ms] org.springframework.web.servlet.HandlerExecutionChain:getHandler() #991
        |       +---[4.248493ms] org.springframework.web.servlet.HandlerAdapter:handle() #991
        |       +---[0.014805ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #993
        |       +---[1.444994ms] org.springframework.web.servlet.DispatcherServlet:applyDefaultViewName() #997
        |       |   `---[0.067631ms] org.springframework.web.servlet.DispatcherServlet:applyDefaultViewName()
        |       +---[0.012027ms] org.springframework.web.servlet.HandlerExecutionChain:applyPostHandle() #998
        |       +---[0.373997ms] org.springframework.web.servlet.DispatcherServlet:processDispatchResult() #1008
        |       |   `---[0.197004ms] org.springframework.web.servlet.DispatcherServlet:processDispatchResult()
        |       |       +---[0.007074ms] org.apache.commons.logging.Log:isDebugEnabled() #1075
        |       |       +---[0.005467ms] org.springframework.web.context.request.async.WebAsyncUtils:getAsyncManager() #1081
        |       |       +---[0.004054ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #1081
        |       |       `---[0.011988ms] org.springframework.web.servlet.HandlerExecutionChain:triggerAfterCompletion() #1087
        |       `---[0.004015ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #1018
        +---[0.005055ms] org.springframework.web.context.request.async.WebAsyncUtils:getAsyncManager() #928
        `---[0.003422ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #928
[jboss@VM_10_91_centos tmp]$ curl -w "@curl-time.txt" http://127.0.0.1:7744/send
success
              http: 200
               dns: 0.001s
          redirect: 0.000s
      time_connect: 0.001s
   time_appconnect: 0.000s
  time_pretransfer: 0.001s
time_starttransfer: 0.115s
     size_download: 7bytes
    speed_download: 60.000B/s
                  ----------
        time_total: 0.115s

本次调用,调用端工夫破费 115 ms,然而从 arthas trace 上看,spring mvc 只耗费了 18ms,那么剩下的 97ms 去哪了呢?

本地测试后曾经能够排除 spring mvc 的问题了,最初也是惟一可能出问题的点就是 tomcat。

可是自己并不相熟 tomcat 中的源码,就连申请入口都不分明,tomcat 里须要 trace 的类都不好找。。。

不过没关系,有神器 Arthas,能够通过 stack 命令来反向查找调用门路,以 org.springframework.web.servlet.DispatcherServlet 作为参数:

stack 输入以后办法被调用的调用门路。

很多时候咱们都晓得一个办法被执行,但这个办法被执行的门路十分多,或者你基本就不晓得这个办法是从那里被执行了,此时你须要的是 stack 命令。

[arthas@24851]$ stack org.springframework.web.servlet.DispatcherServlet *
Press Q or Ctrl+C to abort.
Affect(class-cnt:1 , method-cnt:44) cost in 495 ms.
ts=2019-09-14 21:15:19;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917
    @org.springframework.web.servlet.FrameworkServlet.processRequest()
        at org.springframework.web.servlet.FrameworkServlet.doGet(FrameworkServlet.java:866)
        at javax.servlet.http.HttpServlet.service(HttpServlet.java:635)
        at org.springframework.web.servlet.FrameworkServlet.service(FrameworkServlet.java:851)
        at javax.servlet.http.HttpServlet.service(HttpServlet.java:742)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:231)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.apache.tomcat.websocket.server.WsFilter.doFilter(WsFilter.java:52)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.RequestContextFilter.doFilterInternal(RequestContextFilter.java:99)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.HttpPutFormContentFilter.doFilterInternal(HttpPutFormContentFilter.java:109)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.HiddenHttpMethodFilter.doFilterInternal(HiddenHttpMethodFilter.java:81)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.CharacterEncodingFilter.doFilterInternal(CharacterEncodingFilter.java:200)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.apache.catalina.core.StandardWrapperValve.invoke(StandardWrapperValve.java:198)
        at org.apache.catalina.core.StandardContextValve.invoke(StandardContextValve.java:96)
        at org.apache.catalina.authenticator.AuthenticatorBase.invoke(AuthenticatorBase.java:496)
        at org.apache.catalina.core.StandardHostValve.invoke(StandardHostValve.java:140)
        at org.apache.catalina.valves.ErrorReportValve.invoke(ErrorReportValve.java:81)
        at org.apache.catalina.core.StandardEngineValve.invoke(StandardEngineValve.java:87)
        at org.apache.catalina.connector.CoyoteAdapter.service(CoyoteAdapter.java:342)
        at org.apache.coyote.http11.Http11Processor.service(Http11Processor.java:803)
        at org.apache.coyote.AbstractProcessorLight.process(AbstractProcessorLight.java:66)
        at org.apache.coyote.AbstractProtocol$ConnectionHandler.process(AbstractProtocol.java:790)
        at org.apache.tomcat.util.net.NioEndpoint$SocketProcessor.doRun(NioEndpoint.java:1468)
        at org.apache.tomcat.util.net.SocketProcessorBase.run(SocketProcessorBase.java:49)
        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
        at org.apache.tomcat.util.threads.TaskThread$WrappingRunnable.run(TaskThread.java:61)
        at java.lang.Thread.run(Thread.java:748)
ts=2019-09-14 21:15:19;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917
    @org.springframework.web.servlet.DispatcherServlet.doService()
        at org.springframework.web.servlet.FrameworkServlet.processRequest(FrameworkServlet.java:974)
        at org.springframework.web.servlet.FrameworkServlet.doGet(FrameworkServlet.java:866)
        at javax.servlet.http.HttpServlet.service(HttpServlet.java:635)
        at org.springframework.web.servlet.FrameworkServlet.service(FrameworkServlet.java:851)
        at javax.servlet.http.HttpServlet.service(HttpServlet.java:742)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:231)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.apache.tomcat.websocket.server.WsFilter.doFilter(WsFilter.java:52)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.RequestContextFilter.doFilterInternal(RequestContextFilter.java:99)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.HttpPutFormContentFilter.doFilterInternal(HttpPutFormContentFilter.java:109)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.HiddenHttpMethodFilter.doFilterInternal(HiddenHttpMethodFilter.java:81)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.springframework.web.filter.CharacterEncodingFilter.doFilterInternal(CharacterEncodingFilter.java:200)
        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)
        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)
        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)
        at org.apache.catalina.core.StandardWrapperValve.invoke(StandardWrapperValve.java:198)
        at org.apache.catalina.core.StandardContextValve.invoke(StandardContextValve.java:96)
        at org.apache.catalina.authenticator.AuthenticatorBase.invoke(AuthenticatorBase.java:496)
        at org.apache.catalina.core.StandardHostValve.invoke(StandardHostValve.java:140)
        at org.apache.catalina.valves.ErrorReportValve.invoke(ErrorReportValve.java:81)
        at org.apache.catalina.core.StandardEngineValve.invoke(StandardEngineValve.java:87)
        at org.apache.catalina.connector.CoyoteAdapter.service(CoyoteAdapter.java:342)
        at org.apache.coyote.http11.Http11Processor.service(Http11Processor.java:803)
        at org.apache.coyote.AbstractProcessorLight.process(AbstractProcessorLight.java:66)
        at org.apache.coyote.AbstractProtocol$ConnectionHandler.process(AbstractProtocol.java:790)
        at org.apache.tomcat.util.net.NioEndpoint$SocketProcessor.doRun(NioEndpoint.java:1468)
        at org.apache.tomcat.util.net.SocketProcessorBase.run(SocketProcessorBase.java:49)
        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
        at org.apache.tomcat.util.threads.TaskThread$WrappingRunnable.run(TaskThread.java:61)
        at java.lang.Thread.run(Thread.java:748)

从 stack 日志上能够很直观的看出 DispatchServlet 的调用栈,那么这么长的门路,该 trace 哪个类呢(这里跳过 spring mvc 中的过滤器的 trace 过程,理论排查的时候也 trace 了一遍,但这诡异的工夫耗费不是由这里过滤器产生的)?有肯定教训的老司机从名字上大略也能猜出来从哪里下手比拟好,那就是 org.apache.coyote.http11.Http11Processor.service,从名字上看,http1.1 处理器,这可能是一个比拟好的切入点。上面来 trace 一下:

[arthas@24851]$ trace org.apache.coyote.http11.Http11Processor service
Press Q or Ctrl+C to abort.
Affect(class-cnt:1 , method-cnt:1) cost in 269 ms.
`---ts=2019-09-14 21:22:51;thread_name=http-nio-7744-exec-8;id=17;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418
    `---[131.650285ms] org.apache.coyote.http11.Http11Processor:service()
        +---[0.036851ms] org.apache.coyote.Request:getRequestProcessor() #667
        +---[0.009986ms] org.apache.coyote.RequestInfo:setStage() #668
        +---[0.008928ms] org.apache.coyote.http11.Http11Processor:setSocketWrapper() #671
        +---[0.013236ms] org.apache.coyote.http11.Http11InputBuffer:init() #672
        +---[0.00981ms] org.apache.coyote.http11.Http11OutputBuffer:init() #673
        +---[min=0.00213ms,max=0.007317ms,total=0.009447ms,count=2] org.apache.coyote.http11.Http11Processor:getErrorState() #683
        +---[min=0.002098ms,max=0.008888ms,total=0.010986ms,count=2] org.apache.coyote.ErrorState:isError() #683
        +---[min=0.002448ms,max=0.007149ms,total=0.009597ms,count=2] org.apache.coyote.http11.Http11Processor:isAsync() #683
        +---[min=0.002399ms,max=0.00852ms,total=0.010919ms,count=2] org.apache.tomcat.util.net.AbstractEndpoint:isPaused() #683
        +---[min=0.033587ms,max=0.11832ms,total=0.151907ms,count=2] org.apache.coyote.http11.Http11InputBuffer:parseRequestLine() #687
        +---[0.005384ms] org.apache.tomcat.util.net.AbstractEndpoint:isPaused() #695
        +---[0.007924ms] org.apache.coyote.Request:getMimeHeaders() #702
        +---[0.006744ms] org.apache.tomcat.util.net.AbstractEndpoint:getMaxHeaderCount() #702
        +---[0.012574ms] org.apache.tomcat.util.http.MimeHeaders:setLimit() #702
        +---[0.14319ms] org.apache.coyote.http11.Http11InputBuffer:parseHeaders() #703
        +---[0.003997ms] org.apache.coyote.Request:getMimeHeaders() #743
        +---[0.026561ms] org.apache.tomcat.util.http.MimeHeaders:values() #743
        +---[min=0.002869ms,max=0.01203ms,total=0.014899ms,count=2] java.util.Enumeration:hasMoreElements() #745
        +---[0.070114ms] java.util.Enumeration:nextElement() #746
        +---[0.010921ms] java.lang.String:toLowerCase() #746
        +---[0.008453ms] java.lang.String:contains() #746
        +---[0.002698ms] org.apache.coyote.http11.Http11Processor:getErrorState() #775
        +---[0.00307ms] org.apache.coyote.ErrorState:isError() #775
        +---[0.002708ms] org.apache.coyote.RequestInfo:setStage() #777
        +---[0.171139ms] org.apache.coyote.http11.Http11Processor:prepareRequest() #779
        +---[0.009349ms] org.apache.tomcat.util.net.SocketWrapperBase:decrementKeepAlive() #794
        +---[0.002574ms] org.apache.coyote.http11.Http11Processor:getErrorState() #800
        +---[0.002696ms] org.apache.coyote.ErrorState:isError() #800
        +---[0.002499ms] org.apache.coyote.RequestInfo:setStage() #802
        +---[0.005641ms] org.apache.coyote.http11.Http11Processor:getAdapter() #803
        +---[129.868916ms] org.apache.coyote.Adapter:service() #803
        +---[0.003859ms] org.apache.coyote.http11.Http11Processor:getErrorState() #809
        +---[0.002365ms] org.apache.coyote.ErrorState:isError() #809
        +---[0.003844ms] org.apache.coyote.http11.Http11Processor:isAsync() #809
        +---[0.002382ms] org.apache.coyote.Response:getStatus() #809
        +---[0.002476ms] org.apache.coyote.http11.Http11Processor:statusDropsConnection() #809
        +---[0.002284ms] org.apache.coyote.RequestInfo:setStage() #838
        +---[0.00222ms] org.apache.coyote.http11.Http11Processor:isAsync() #839
        +---[0.037873ms] org.apache.coyote.http11.Http11Processor:endRequest() #843
        +---[0.002188ms] org.apache.coyote.RequestInfo:setStage() #845
        +---[0.002112ms] org.apache.coyote.http11.Http11Processor:getErrorState() #849
        +---[0.002063ms] org.apache.coyote.ErrorState:isError() #849
        +---[0.002504ms] org.apache.coyote.http11.Http11Processor:isAsync() #853
        +---[0.009808ms] org.apache.coyote.Request:updateCounters() #854
        +---[0.002008ms] org.apache.coyote.http11.Http11Processor:getErrorState() #855
        +---[0.002192ms] org.apache.coyote.ErrorState:isIoAllowed() #855
        +---[0.01968ms] org.apache.coyote.http11.Http11InputBuffer:nextRequest() #856
        +---[0.010065ms] org.apache.coyote.http11.Http11OutputBuffer:nextRequest() #857
        +---[0.002576ms] org.apache.coyote.RequestInfo:setStage() #870
        +---[0.016599ms] org.apache.coyote.http11.Http11Processor:processSendfile() #872
        +---[0.008182ms] org.apache.coyote.http11.Http11InputBuffer:getParsingRequestLinePhase() #688
        +---[0.0075ms] org.apache.coyote.http11.Http11Processor:handleIncompleteRequestLineRead() #690
        +---[0.001979ms] org.apache.coyote.RequestInfo:setStage() #875
        +---[0.001981ms] org.apache.coyote.http11.Http11Processor:getErrorState() #877
        +---[0.001934ms] org.apache.coyote.ErrorState:isError() #877
        +---[0.001995ms] org.apache.tomcat.util.net.AbstractEndpoint:isPaused() #877
        +---[0.002403ms] org.apache.coyote.http11.Http11Processor:isAsync() #879
        `---[0.006176ms] org.apache.coyote.http11.Http11Processor:isUpgrade() #881

日志里有一个 129ms 的耗时点(工夫比没开 arthas 的时候更长是因为 arthas 自身带来的性能耗费,所以生产环境小心应用),这个就是要找的问题点。

打问题点找到了,那怎么定位是什么导致的问题呢,又如何解决呢?

持续 trace 吧,细化到具体的代码块或者内容。trace 因为性能思考,不会展现所有的调用门路,如果调用门路过深,只有手动深刻 trace,准则就是 trace 耗时长的那个办法:

[arthas@24851]$ trace org.apache.coyote.Adapter service
Press Q or Ctrl+C to abort.
Affect(class-cnt:1 , method-cnt:1) cost in 608 ms.
`---ts=2019-09-14 21:34:33;thread_name=http-nio-7744-exec-1;id=10;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418
    `---[81.70999ms] org.apache.catalina.connector.CoyoteAdapter:service()
        +---[0.032546ms] org.apache.coyote.Request:getNote() #302
        +---[0.007148ms] org.apache.coyote.Response:getNote() #303
        +---[0.007475ms] org.apache.catalina.connector.Connector:getXpoweredBy() #324
        +---[0.00447ms] org.apache.coyote.Request:getRequestProcessor() #331
        +---[0.007902ms] java.lang.ThreadLocal:get() #331
        +---[0.006522ms] org.apache.coyote.RequestInfo:setWorkerThreadName() #331
        +---[73.793798ms] org.apache.catalina.connector.CoyoteAdapter:postParseRequest() #336
        +---[0.001536ms] org.apache.catalina.connector.Connector:getService() #339
        +---[0.004469ms] org.apache.catalina.Service:getContainer() #339
        +---[0.007074ms] org.apache.catalina.Engine:getPipeline() #339
        +---[0.004334ms] org.apache.catalina.Pipeline:isAsyncSupported() #339
        +---[0.002466ms] org.apache.catalina.connector.Request:setAsyncSupported() #339
        +---[6.01E-4ms] org.apache.catalina.connector.Connector:getService() #342
        +---[0.001859ms] org.apache.catalina.Service:getContainer() #342
        +---[9.65E-4ms] org.apache.catalina.Engine:getPipeline() #342
        +---[0.005231ms] org.apache.catalina.Pipeline:getFirst() #342
        +---[7.239154ms] org.apache.catalina.Valve:invoke() #342
        +---[0.006904ms] org.apache.catalina.connector.Request:isAsync() #345
        +---[0.00509ms] org.apache.catalina.connector.Request:finishRequest() #372
        +---[0.051461ms] org.apache.catalina.connector.Response:finishResponse() #373
        +---[0.007244ms] java.util.concurrent.atomic.AtomicBoolean:<init>() #379
        +---[0.007314ms] org.apache.coyote.Response:action() #380
        +---[0.004518ms] org.apache.catalina.connector.Request:isAsyncCompleting() #382
        +---[0.001072ms] org.apache.catalina.connector.Request:getContext() #394
        +---[0.007166ms] java.lang.System:currentTimeMillis() #401
        +---[0.004367ms] org.apache.coyote.Request:getStartTime() #401
        +---[0.011483ms] org.apache.catalina.Context:logAccess() #401
        +---[0.0014ms] org.apache.coyote.Request:getRequestProcessor() #406
        +---[min=8.0E-4ms,max=9.22E-4ms,total=0.001722ms,count=2] java.lang.Integer:<init>() #406
        +---[0.001082ms] java.lang.reflect.Method:invoke() #406
        +---[0.001851ms] org.apache.coyote.RequestInfo:setWorkerThreadName() #406
        +---[0.035805ms] org.apache.catalina.connector.Request:recycle() #410
        `---[0.007849ms] org.apache.catalina.connector.Response:recycle() #411

一段无聊的手动深刻 trace 之后………………

[arthas@24851]$ trace org.apache.catalina.webresources.AbstractArchiveResourceSet getArchiveEntries
Press Q or Ctrl+C to abort.
Affect(class-cnt:4 , method-cnt:2) cost in 150 ms.
`---ts=2019-09-14 21:36:26;thread_name=http-nio-7744-exec-3;id=12;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418
    `---[75.743681ms] org.apache.catalina.webresources.JarWarResourceSet:getArchiveEntries()
        +---[0.025731ms] java.util.HashMap:<init>() #106
        +---[0.097729ms] org.apache.catalina.webresources.JarWarResourceSet:openJarFile() #109
        +---[0.091037ms] java.util.jar.JarFile:getJarEntry() #110
        +---[0.096325ms] java.util.jar.JarFile:getInputStream() #111
        +---[0.451916ms] org.apache.catalina.webresources.TomcatJarInputStream:<init>() #113
        +---[min=0.001175ms,max=0.001176ms,total=0.002351ms,count=2] java.lang.Integer:<init>() #114
        +---[0.00104ms] java.lang.reflect.Method:invoke() #114
        +---[0.045105ms] org.apache.catalina.webresources.TomcatJarInputStream:getNextJarEntry() #114
        +---[min=5.02E-4ms,max=0.008531ms,total=0.028864ms,count=31] java.util.jar.JarEntry:getName() #116
        +---[min=5.39E-4ms,max=0.022805ms,total=0.054647ms,count=31] java.util.HashMap:put() #116
        +---[min=0.004452ms,max=34.479307ms,total=74.206249ms,count=31] org.apache.catalina.webresources.TomcatJarInputStream:getNextJarEntry() #117
        +---[0.018358ms] org.apache.catalina.webresources.TomcatJarInputStream:getManifest() #119
        +---[0.006429ms] org.apache.catalina.webresources.JarWarResourceSet:setManifest() #120
        +---[0.010904ms] org.apache.tomcat.util.compat.JreCompat:isJre9Available() #121
        +---[0.003307ms] org.apache.catalina.webresources.TomcatJarInputStream:getMetaInfEntry() #133
        +---[5.5E-4ms] java.util.jar.JarEntry:getName() #135
        +---[6.42E-4ms] java.util.HashMap:put() #135
        +---[0.001981ms] org.apache.catalina.webresources.TomcatJarInputStream:getManifestEntry() #137
        +---[0.064484ms] org.apache.catalina.webresources.TomcatJarInputStream:close() #141
        +---[0.007961ms] org.apache.catalina.webresources.JarWarResourceSet:closeJarFile() #151
        `---[0.004643ms] java.io.InputStream:close() #155

发现了一个值得暂停思考的点:

+---[min=0.004452ms,max=34.479307ms,total=74.206249ms,count=31] org.apache.catalina.webresources.TomcatJarInputStream:getNextJarEntry() #117

这行代码加载了 31 次,一共耗时 74ms;从名字上看,应该是 tomcat 加载 jar 包时的耗时,那么是加载了 31 个 jar 包的耗时,还是加载了 jar 包内的某些资源 31 次耗时呢?

TomcatJarInputStream 这个类源码的正文写到:

The purpose of this sub-class is to obtain references to the JarEntry objectsfor META-INF/ and META-INF/MANIFEST.MF that are otherwise swallowed by theJarInputStream implementation.

大略意思也就是,获取 jar 包内 META-INF/,META-INF/MANIFEST 的资源,这是一个子类,更多的性能在父类 JarInputStream 里。

其实看到这里大略也能猜到问题了,tomcat 加载 jar 包内 META-INF/,META-INF/MANIFEST 的资源导致的耗时,至于为什么间断申请不会耗时,应该是 tomcat 的缓存机制(上面介绍源码剖析)。

不焦急定位问题,试着通过 Arthas 最终定位问题细节,持续手动深刻 trace。

[arthas@24851]$ trace org.apache.catalina.webresources.TomcatJarInputStream *
Press Q or Ctrl+C to abort.
Affect(class-cnt:1 , method-cnt:4) cost in 44 ms.
`---ts=2019-09-14 21:37:47;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418
    `---[0.234952ms] org.apache.catalina.webresources.TomcatJarInputStream:createZipEntry()
        +---[0.039455ms] java.util.jar.JarInputStream:createZipEntry() #43
        `---[0.007827ms] java.lang.String:equals() #44
`---ts=2019-09-14 21:37:47;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418
    `---[0.050222ms] org.apache.catalina.webresources.TomcatJarInputStream:createZipEntry()
        +---[0.001889ms] java.util.jar.JarInputStream:createZipEntry() #43
        `---[0.001643ms] java.lang.String:equals() #46
#这里一共 31 个 trace 日志,删减了剩下的

从办法名上看,还是加载资源之类的意思。都曾经到 jdk 源码了,这时候来看一下 TomcatJarInputStream 这个类的源码:

/**
 * Creates a new <code>JarEntry</code> (<code>ZipEntry</code>) for the
 * specified JAR file entry name. The manifest attributes of
 * the specified JAR file entry name will be copied to the new
 * <CODE>JarEntry</CODE>.
 *
 * @param name the name of the JAR/ZIP file entry
 * @return the <code>JarEntry</code> object just created
 */
protected ZipEntry createZipEntry(String name) {JarEntry e = new JarEntry(name);
    if (man != null) {e.attr = man.getAttributes(name);
    }
    return e;
}

这个 createZipEntry 有个 name 参数,从正文上看,是 jar/zip 文件名,如果能失去文件名这种要害信息,就能够间接定位问题了;还是通过 Arthas,应用 watch 命令,动静监测办法调用数据。

watch 办法执行数据观测

让你能不便的察看到指定办法的调用状况。能察看到的范畴为:返回值、抛出异样、入参,通过编写 OGNL 表达式进行对应变量的查看。

watch 该办法的入参:

[arthas@24851]$ watch  org.apache.catalina.webresources.TomcatJarInputStream createZipEntry "{params[0]}"
Press Q or Ctrl+C to abort.
Affect(class-cnt:1 , method-cnt:1) cost in 27 ms.
ts=2019-09-14 21:51:14; [cost=0.14547ms] result=@ArrayList[@String[META-INF/],
]
ts=2019-09-14 21:51:14; [cost=0.048028ms] result=@ArrayList[@String[META-INF/MANIFEST.MF],
]
ts=2019-09-14 21:51:14; [cost=0.046071ms] result=@ArrayList[@String[META-INF/resources/],
]
ts=2019-09-14 21:51:14; [cost=0.033855ms] result=@ArrayList[@String[META-INF/resources/swagger-ui.html],
]
ts=2019-09-14 21:51:14; [cost=0.039138ms] result=@ArrayList[@String[META-INF/resources/webjars/],
]
ts=2019-09-14 21:51:14; [cost=0.033701ms] result=@ArrayList[@String[META-INF/resources/webjars/springfox-swagger-ui/],
]
ts=2019-09-14 21:51:14; [cost=0.033644ms] result=@ArrayList[@String[META-INF/resources/webjars/springfox-swagger-ui/favicon-16x16.png],
]
ts=2019-09-14 21:51:14; [cost=0.033976ms] result=@ArrayList[@String[META-INF/resources/webjars/springfox-swagger-ui/springfox.css],
]
ts=2019-09-14 21:51:14; [cost=0.032818ms] result=@ArrayList[@String[META-INF/resources/webjars/springfox-swagger-ui/swagger-ui-standalone-preset.js.map],
]
ts=2019-09-14 21:51:14; [cost=0.04651ms] result=@ArrayList[@String[META-INF/resources/webjars/springfox-swagger-ui/swagger-ui.css],
]
ts=2019-09-14 21:51:14; [cost=0.034793ms] result=@ArrayList[@String[META-INF/resources/webjars/springfox-swagger-ui/swagger-ui.js.map],

这下间接看到了具体加载的资源名,这么相熟的名字:swagger-ui,一个国外的 rest 接口文档工具,又有国内开发者基于 swagger-ui 做了一套 spring mvc 的集成工具,通过注解就能够主动生成 swagger-ui 须要的接口定义 json 文件,用起来还比拟不便,就是侵入性较强。

删除 swagger 的 jar 包后问题,诡异的 70+ms 就隐没了。

<!--pom 里删除这两个援用,这两个包时国内开发者封装的,swagger-ui 并没有提供 java spring-mvc 的反对包,swagger 只是一个浏览器端的 ui+editor
<dependency>
    <groupId>io.springfox</groupId>
    <artifactId>springfox-swagger2</artifactId>
    <version>2.9.2</version>
</dependency>
<dependency>
    <groupId>io.springfox</groupId>
    <artifactId>springfox-swagger-ui</artifactId>
    <version>2.9.2</version>
</dependency>

那么为什么 swagger 会导致申请耗时呢,为什么每次申请偶读会加载 swagger 外部的动态资源呢?

其实这是 tomcat-embed 的一个 bug 吧,上面具体介绍一下该 Bug。

Tomcat embed Bug 剖析 & 解决

源码剖析过程切实太漫长,而且也不是本文的重点,所以就不介绍了,上面间接介绍下剖析后果。

顺便贴一张 tomcat 解决申请的外围类图:

1. 为什么每次申请会加载 Jar 包内的动态资源?

关键在于 org.apache.catalina.mapper.Mapper#internalMapWrapper 这个办法,该版本下解决申请的形式有问题,导致每次都校验动态资源。

2. 为什么间断申请不会呈现问题?

因为 Tomcat 对于这种动态资源的解析是有缓存的,优先从缓存查找,缓存过期后再从新解析。具体参考 org.apache.catalina.webresources.Cache,默认过期工夫 ttl 是 5000ms。

3. 为什么本地不会复现?

其实确切的说,是通过 spring-boot 打包插件后不能复现。因为启动形式的不同,tomcat 应用了不同的类去解决动态资源,所以没问题。

4. 如何解决?

1)降级 tomcat-embed 版本即可

以后呈现 Bug 的版本为:spring-boot:2.0.2.RELEASE,内置的 tomcat embed 版本为 8.5.31。降级 tomcat embed 版本至 8.5.40+ 即可解决此问题,新版本曾经修复了。

2)通过替换 springboot pom properties 形式

如果我的项目是 maven 是继承的 springboot,即 parent 配置为 springboot 的,或者 dependencyManagement 中 import spring boot 包的。

<parent>
        <groupId>org.springframework.boot</groupId>
        <artifactId>spring-boot-starter-parent</artifactId>
        <version>2.0.2.RELEASE</version>
        <relativePath/> <!-- lookup parent from repository -->
    </parent>

pom 中间接笼罩 properties 即可:

<properties>
    <tomcat.version>8.5.40</tomcat.version>
</properties>

3)降级 spring boot 版本

springboot 2.1.0.RELEASE 中的 tomcat embed 版本曾经大于 8.5.31 了,所以间接将 springboot 降级至该版本及以上版本就能够解决此问题。

作者简介

空无,Arthas 铁粉,一个酷爱技术酷爱分享的程序员,专一 JAVA 后端开发。

欢送登陆 start.aliyun.com 知口头手实验室体验 Arthas 57 个入手试验:https://start.aliyun.com/handson-lab/#!category=arthas

为了让更多开发者开始用上 Arthas 这个 Java 诊断神器,Arthas 社区联结 JetBrains 推出 Arthas 有奖征文活动:聊聊这些年你和 Arthas 之间的那些事儿。流动仍在炽热进行中,点击即可参加,欢送大家踊跃投稿,参加即有可能获奖!

正文完
 0